Anscombe's Quartet

The cover image about Anscombe’s Quartet was generated by ChatGPT, and it used the following prompt: “The digital design highlights the title ‘Anscombe’s Quartet’ in bold, white sans-serif letters, centered against a dynamic abstract backdrop. The image is split into four colorful quadrants, each showcasing unique textures and patterns—ranging from painterly hues and curved lines to scattered circles and dots.”
Introduction
While listening to a presentation today, I happened to hear a term——Anscombe’s Quartet, a term I had never encountered before, yet it has significant implications in both statistics and data visualization.
History
Anscombe’s quartet consists of four datasets constructed by British statistician Francis Anscombe in 1973. These four datasets share nearly identical statistical properties but differ dramatically when visualized graphically.
Anscombe’s Quartet
Here are the four sets of data, each containing 11 pairs of $x$ and $y$ values:
| $x_1$ | $y_1$ | $x_2$ | $y_2$ | $x_3$ | $y_3$ | $x_4$ | $y_4$ | |||
|---|---|---|---|---|---|---|---|---|---|---|
| 10.0 | 8.04 | 10.0 | 9.14 | 10.0 | 7.46 | 8.0 | 6.58 | |||
| 8.0 | 6.95 | 8.0 | 8.14 | 8.0 | 6.77 | 8.0 | 5.76 | |||
| 13.0 | 7.58 | 13.0 | 8.74 | 13.0 | 12.74 | 8.0 | 7.71 | |||
| 9.0 | 8.81 | 9.0 | 8.77 | 9.0 | 7.11 | 8.0 | 8.84 | |||
| 11.0 | 8.33 | 11.0 | 9.26 | 11.0 | 7.81 | 8.0 | 8.47 | |||
| 14.0 | 9.96 | 14.0 | 8.10 | 14.0 | 8.84 | 8.0 | 7.04 | |||
| 6.0 | 7.24 | 6.0 | 6.13 | 6.0 | 6.08 | 8.0 | 5.25 | |||
| 4.0 | 4.26 | 4.0 | 3.10 | 4.0 | 5.39 | 19.0 | 12.50 | |||
| 12.0 | 10.84 | 12.0 | 9.13 | 12.0 | 8.15 | 8.0 | 5.56 | |||
| 7.0 | 4.82 | 7.0 | 7.26 | 7.0 | 6.42 | 8.0 | 7.91 | |||
| 5.0 | 5.68 | 5.0 | 4.74 | 5.0 | 5.73 | 8.0 | 6.89 |
We’ll now analyze this dataset using the R programming language.
Data Input
We enter the data as matrices in R for ease of manipulation:
| |
| |
Mean, Variance, and Correlation Coefficient
We examine the mean, variance, and correlation coefficients of the data.
Interestingly, the means and variances are nearly identical up to two decimal places. The correlation coefficients for each dataset pair, such as $(x_1, y_1)$, are also nearly identical.
| |
| |
Linear Regression
We apply linear regression to each dataset. All four datasets yield similar regression equations:
$$ y = 3 + 0.5 x. $$
| |
| |
This is quite astonishing! Without further checks, one might assume the datasets are essentially the same, but they are not.
Scatter Plots
Here are the scatter plots for the four datasets. They tell a different story.
Dataset 1 resembles a linear relationship; Dataset 2 shows a clear non-linear trend; Dataset 3 includes an outlier that affects the regression; Dataset 4 consists mostly of constant $x$ values with a single outlier.
| |

Boxplots
Boxplots show that the quartile distributions of each dataset vary as well.
| |

Conclusion
Anscombe’s Quartet teaches us a critical lesson: we should not rely solely on statistical summaries. Instead, we must leverage tools like data visualization and diverse analytical perspectives to fully understand our data. Statistical figures are essential, but without proper context and visualization, they can be misleading.
Further Learning
- The R Notebook HTML file used in this article.
References
- Anscombe’s Quartet. (September 23, 2021). Wikipedia, The Free Encyclopedia. Retrieved June 3, 2025 from https://zh.wikipedia.org/zh-tw/安斯库姆四重奏

![[Thought] Historical Earthquake Locations Around Taiwan](https://Josh-test-lab.github.io/posts/Historical%20Earthquake%20Locations%20Around%20Taiwan/cover%20image.webp)






