What a scatter plot shows
A scatter plot puts one point on the page for each observation in your data. Each point carries two measurements: its left-to-right position is the x value, and its up-and-down position is the y value. You use it when you want to see whether two things move together. If high x values tend to sit with high y values, the cloud of points slopes up. If they tend to sit with low y values, it slopes down. If there is no pattern, the points look like a loose swarm.
This page is the scatter plot maker at AxisPlot. Paste your numbers, and the chart appears with axes, a fitted straight line and R squared. Nothing is uploaded. The code runs in your browser.
The exact shape of data you need
Two numeric columns. The x column first, the y column second. One row per observation. That is the whole requirement.
Here is a worked example. Suppose you paste these six rows:
| x | y |
|---|---|
| 1 | 2.1 |
| 2 | 3.9 |
| 3 | 6.2 |
| 4 | 7.8 |
| 5 | 10.1 |
| 6 | 11.9 |
The tool reads the first row as headings because the rows below it are numbers and the first row is not. It fits an ordinary least squares line, which is the straight line with the smallest sum of squared vertical distances to the points. For these six points the line is close to y = 0.5 + 1.95x, and R squared is about 0.998. That means the straight line accounts for almost all of the up-and-down variation in y.
If your first row is plain text and the rows below are also plain text, the tool keeps that first row as data. It only treats a row as headings when the rows below it are numbers or dates and the first row is not.
Rows with fewer columns than the widest row are padded, and the count is reported. They are never dropped. Empty cells and spreadsheet placeholders such as n/a, NA, NaN, null, a lone dash, #DIV/0! and #VALUE! are counted as empty rather than treated as zero. That matters because one n/a should not turn a column of numbers into a column of text.
How to read the result
Look at four things: direction, strength, clusters and outliers.
- Direction. Does the cloud slope up, slope down, or lie flat?
- Strength. How tightly do the points hug a line? A tight cigar is strong. A wide blob is weak.
- Clusters. Are there groups of points sitting apart from the rest? A cluster can mean two different populations are mixed in one column.
- Outliers. Is one point far from the others? A single outlier can pull the fitted line toward itself.
The trend line is a summary, not a cause. It says what a straight line through these points would look like. It does not say that x makes y change.
R squared is the share of the up-and-down in y that the straight line accounts for. It runs from 0 to 1. A value near 1 means the line tracks the points closely. A value near 0 means the line does not help much. R squared says nothing about whether a straight line was the right shape. Anscombe's four datasets have the same means, the same variances, the same correlation of 0.816 and the same fitted line y = 3 + 0.5x, yet one is a clean straight relationship, one is a curve, one is a straight line with a single outlier and one is a vertical stack of points with one point far away [Anscombe 1973]. Looking at the points is not optional.
Correlation is not causation. Two things can move together because a third thing drives both, or by coincidence. A chart shows what is in the numbers you pasted, and no more.
A vertical set of points has no least-squares fit. The page says so rather than drawing a line.
Log scale, skipped rows and when to use another chart
A logarithmic vertical axis is offered for scatter plots and is only applied when every value is above zero. It helps when your y values span several orders of magnitude, such as 1, 10, 100 and 1000. On a log scale, equal distances mean equal ratios rather than equal amounts. That can turn a curve into a straight line, which makes the pattern easier to see.
Non-numeric rows are skipped and counted. The page tells you how many rows were skipped so you can check whether the right data went in. Missing values are treated as empty, not as zero.
Sometimes a scatter plot is the wrong choice. If you are comparing amounts across categories, use a bar chart. If something changes over time or along another ordered axis, use a line chart, because a line asserts that neighbouring points are connected, which is false for unordered categories. If you want the shape of a single column of numbers, use a histogram. If you are showing parts of one whole and only a handful of parts, use a pie chart. The chart maker on the home page can point you to the right one.
You get a PNG drawn at twice the chart's size so it stays sharp in a document or a slide, and an SVG that stays sharp at any size and can be opened in Illustrator, Inkscape, Figma or Word. There is no watermark, no sign-up and no limit on the number of charts. The chart carries a title and a description in the file itself so a screen reader can announce it, and the figures behind every chart are also shown as a table on the page, because a chart is an image and an image needs a text equivalent [WCAG 2.2, success criterion 1.1.1].
The maths behind the fitted line, the axis ticks and the summary statistics is on how it works. That page also lists the full references for every source named here.