What Is PC1 and PC2 in Principal Component Analysis?

PC1 and PC2 are the first and second principal components produced by principal component analysis, a technique that distills a sprawling dataset with many variables into a smaller set of new, synthetic variables ranked by how much of the data’s overall variation each one captures. PC1 is the direction through the data along which values spread out the most, and PC2 is the direction of next-greatest spread, constrained to be perpendicular to PC1. Together they form a new coordinate system that often captures the lion’s share of what is happening in a dataset, which is why researchers routinely plot PC1 against PC2 as a first look at their data’s structure.

Why We Need New Axes in the First Place

Imagine you have measured dozens of things about each item in your dataset: gene expression levels across thousands of genes, or dozens of chemical concentrations in a set of biological samples. Many of those measurements move together. When one goes up, another tends to follow. That overlap means you are storing the same underlying signal in multiple columns. PCA creates new variables that successively maximize variance while being completely uncorrelated with one another, effectively stripping out that redundancy and boiling the dataset down to its most informative dimensions.1PubMed Central. Principal component analysis: a review and recent developments The result is a set of principal components, each accounting for a progressively smaller slice of the total variation, ordered from PC1 (the biggest slice) on down.

The word “component” can be misleading. A principal component is not one of the original measurements pulled out and relabeled. It is a weighted combination of all the original variables, chosen so that the resulting composite captures the maximum possible spread in the data. PC1 is the single best summary of the dataset in one dimension. PC2 is the best summary you can add after PC1 has already been accounted for. PC3, PC4, and so on continue the process, each one smaller in importance than the last.

How PC1 and PC2 Are Calculated

At its core, PCA works by finding the directions through the data cloud along which the data points are most spread out. Picture a swarm of dots in space. PC1 is the line you could draw through that swarm so that, if you projected every dot onto the line, the projected points would be as spread apart as possible. PC2 is the next line you draw, forced to be at a right angle to PC1, that captures the most remaining spread.1PubMed Central. Principal component analysis: a review and recent developments This right-angle constraint is what statisticians call orthogonality, and it guarantees that PC1 and PC2 carry no overlapping information.

The math behind it involves finding eigenvectors of the data’s covariance matrix, but the intuition matters more than the algebra. Each eigenvector points in a direction of maximal variance, and the associated eigenvalue tells you how much variance that direction accounts for.1PubMed Central. Principal component analysis: a review and recent developments The eigenvector with the largest eigenvalue becomes PC1, the next largest becomes PC2, and so on. If your dataset has 50 original variables, PCA produces 50 principal components, but the hope is that the first two or three explain most of the variation and the rest can be safely ignored.

Scores and Loadings, the Two Things People Actually Look At

When someone says “plot PC1 vs. PC2,” they almost always mean a scores plot. Each data point (a sample, a person, a country) gets a new pair of coordinates: its position along the PC1 axis and its position along the PC2 axis. These coordinates are the scores, and they tell you where that sample sits in the simplified, two-dimensional summary of the data.2PubMed. Exploration of Principal Component Analysis: Deriving Principal Component Analysis Visually Using Spectra Clusters of dots on a scores plot reveal groups of samples that behave similarly across the original measurements. Outliers stand apart, and gradients suggest continuous variation.

Loadings answer a different question: which of the original variables contribute most to each principal component? Every original variable gets a loading value for PC1 and a separate loading value for PC2. A high loading on PC1 means that variable is strongly aligned with the direction of greatest variation in the data. Looking at loadings is how researchers move from “these samples form two clusters” to “the difference between those clusters is mostly driven by variables X, Y, and Z.” Each principal component is essentially a weighted recipe of original variables, and the loadings are the recipe’s ingredient weights.2PubMed. Exploration of Principal Component Analysis: Deriving Principal Component Analysis Visually Using Spectra

What the Percentage of Variance Explained Actually Means

You will almost always see a label on PCA plots like “PC1 (45%)” and “PC2 (18%).” Those percentages tell you the fraction of total variation in the entire dataset that each component captures. If PC1 explains 45 percent of the variance, that single synthetic axis carries nearly half of all the information spread across all the original variables. If PC1 and PC2 together explain 63 percent, you know the two-dimensional plot you are looking at is a decent, though incomplete, picture of the full dataset.

There is no universal threshold for “good enough.” In some fields, two components capturing 80 or 90 percent of the variance is routine and means the data had a strong underlying structure. In messier domains like social science surveys or ecological community data, PC1 might only explain 15 percent of the total variance, and you might need many components before you have accounted for most of what is going on. A low percentage does not mean PCA failed; it means the data’s variation is spread across many independent dimensions rather than concentrated in a few.

A common mistake is interpreting a high PC1 percentage as proof that one factor “controls” the data. PC1 is the direction of maximum spread, but that spread could arise from a single powerful underlying cause or from many correlated causes that happen to push in the same direction. PCA identifies patterns, not causes. Two datasets with very different causal structures can produce nearly identical PC1 axes.

Why the PC1 vs. PC2 Plot Is So Popular

Humans see well in two dimensions. A scatter plot of PC1 against PC2 is the single most informative flat picture you can draw of a high-dimensional dataset, because PC1 and PC2 together capture more variation than any other pair of axes. It is the go-to exploratory visualization across nearly every quantitative discipline: genomics, chemistry, ecology, economics, psychology. Researchers use it to spot clusters, outliers, and gradients before running more formal analyses.

That said, collapsing dozens or hundreds of dimensions into two inevitably loses information. Real structure can hide in PC3, PC4, or beyond. A pair of groups that overlap on the PC1-vs.-PC2 plot may separate cleanly on PC3. Experienced analysts typically inspect several pairwise plots (PC1 vs. PC3, PC2 vs. PC3, and so on) before drawing conclusions, and they check the variance-explained percentages to judge how much they might be missing.

PCA in Genetics and Ancestry Studies

One of the most recognizable uses of PC1 and PC2 is in population genetics. Researchers feed hundreds of thousands of genetic variants into PCA, and the resulting scores plot often neatly separates individuals by continental ancestry. PC1 of genetic data routinely used to infer ancestry and control for population structure in genetic analyses.3Bioinformatics. Efficient toolkit implementing best practices for principal component analysis of population genetic data In large biobank projects, researchers plot PC1 versus PC2 alongside reference populations from around the world, with the percent of variance explained by each component labeled on the axes.4Nature Communications. Genetic ancestry and population structure in the All of Us Research Program cohort

In these studies, PC1 often separates broadly along one major axis of global genetic diversity (for instance, African from non-African ancestry), while PC2 captures the next-largest axis of differentiation (often European versus East Asian ancestry, depending on the dataset). Additional PCs can tease apart finer-grained structure, such as regional variation within a continent. The reason PCA works so well here is that human migration history has created correlated patterns across many genetic markers, and PCA excels at detecting exactly those kinds of coordinated shifts.

Geneticists also include the first several PCs as covariates in statistical models to avoid false associations. If a disease happens to be more common in one ancestry group, and a genetic variant is also more common in that group for purely historical reasons, a naive analysis might mistake the variant for a disease gene. Adjusting for PC1, PC2, and a handful of further components helps control for this confounding.

PCA in Psychology and the Concept of a General Factor

In psychometrics, batteries of cognitive tests tend to be positively correlated: people who score well on vocabulary also tend to score well on spatial reasoning and processing speed, for example. When researchers run PCA on a set of cognitive test scores, PC1 captures that shared positive correlation, and it is sometimes interpreted as a general factor of cognitive ability, often labeled “g.” In one study using 19 psychometric variables, the general factor was represented by PC1.5Intelligence. Occupation and income related to psychometric g PC2 and subsequent components then pick up the ways specific tests differ from that shared pattern, such as contrasts between verbal and spatial ability.

This is a good illustration of what PC1 does and does not tell you. It reliably extracts the dominant shared pattern across variables. But whether you call that pattern “general intelligence,” “test familiarity,” or “socioeconomic advantage” is an interpretive choice, not something PCA resolves. PCA describes the shape of the variation; naming the cause is a separate scientific argument.

PCA in Face Recognition and Image Compression

A grayscale photograph can be thought of as a list of pixel intensities, one number per pixel. A face image that is 100 by 100 pixels has 10,000 variables. Running PCA on a collection of face images produces principal components that are themselves face-shaped images, sometimes called eigenfaces. PC1 captures the most common pattern of brightness variation across all the faces, often related to overall illumination. PC2 and subsequent components pick up progressively subtler variations in face shape, expression, and lighting.6ScienceDirect. Face recognition based on PCA image reconstruction and LDA

The practical payoff is enormous. Instead of comparing two images pixel by pixel across all 10,000 dimensions, you compare their scores on, say, the first 50 principal components. This dimensionality reduction makes face-recognition algorithms faster and often more accurate, because the first few PCs tend to capture meaningful structural differences while the later ones mostly encode noise. The idea was introduced in the early 1990s and remains a foundation of many image processing pipelines, even as deep-learning approaches have become dominant.

What PCA Assumes and Where It Breaks Down

PCA assumes that the important patterns in the data lie along straight-line directions. If two groups of samples separate along a curve or a spiral rather than along a straight axis, PCA can miss the structure entirely or split it awkwardly across many components. The linearity of PCA limits its power for complex datasets because it cannot capture relationships defined by anything beyond simple correlations.7ScienceDirect. Adaptive nonlinear manifolds and their applications to pattern recognition Techniques like kernel PCA, t-SNE, and UMAP were developed specifically to handle curved or clustered structures that linear PCA flattens.

PCA is also sensitive to the scale of the original variables. If one variable is measured in meters and another in millimeters, the millimeter variable will dominate the variance simply because its numbers are larger. For this reason, most implementations standardize the data before running PCA, centering each variable to zero mean and scaling it to unit variance. Without this step, PC1 may just reflect whichever variable happened to have the largest raw numbers, which is rarely the pattern you want to find.

Outliers can distort results as well. Because PCA maximizes variance, a handful of extreme data points can pull PC1 toward themselves, warping the entire coordinate system. Checking for and handling outliers before running PCA is standard practice in most fields.

Rotation and the Interpretability Problem

Raw principal components can be hard to interpret because each one is a mixture of all original variables. PC1 might have moderate loadings on 30 different genes, none of them obviously dominant. Rotation methods, the most common being varimax rotation, redistribute the variance among the components so that each rotated component loads heavily on a small number of variables and near zero on the rest.8Mathematics. New Modeling Approaches Based on Varimax Rotation of Functional Principal Components This makes each component easier to name and understand, at the cost of breaking the strict variance-ordering property: after rotation, the first component no longer necessarily explains the most variance.

Rotation does not change the total amount of variance explained by the set of retained components. It simply redistributes it. Think of it as rearranging furniture in a room without adding or removing anything. Whether to rotate depends on the goal. If the aim is pure dimensionality reduction or visualization, unrotated PCs are usually fine. If the aim is to identify interpretable latent factors, rotation is the norm, particularly in survey research and psychometrics.

PC1 and PC2 Do Not Always Mean the Same Thing Across Studies

A subtle but important point that trips up newcomers: the meaning of PC1 is entirely determined by the data fed into the analysis. In a genetics study, PC1 might reflect continental ancestry. In a consumer survey, PC1 might reflect overall spending level. In a spectroscopy dataset, PC1 might reflect baseline differences in sample thickness. There is no universal “first component” that carries a fixed interpretation.

Even within the same field, adding or removing variables, or adding or removing samples, can change what PC1 and PC2 represent. If you run PCA on gene expression from blood samples and then rerun it after adding brain tissue samples, the principal components will shift because the dominant patterns of variation have changed. This means you cannot directly compare PC1 scores from two separate PCA runs unless the analyses used the same variables measured on similar populations.

Researchers sometimes project new samples onto an existing PCA model, which preserves the original component definitions. This is common in genetics, where a reference panel defines the PC axes and new individuals are placed onto those axes. That approach keeps PC1 and PC2 interpretable in the same terms across datasets, but it requires deliberate effort.

Common Misreadings of a PC1-vs.-PC2 Plot

The most frequent mistake is reading distance on the plot as a meaningful similarity measure without checking the variance-explained percentages. If PC1 explains 40 percent of the variance and PC2 explains 5 percent, two samples that are far apart along PC2 but close along PC1 are actually much more similar than they appear on the plot. The axes are not on equal footing, and the raw scatter plot can exaggerate differences along the minor axis.

Another common error is treating clusters on the plot as definitive groups. PCA is a visualization and dimensionality-reduction tool, not a clustering algorithm. Apparent clusters can sometimes be artifacts of projecting continuous variation onto two dimensions, and distinct groups can sometimes overlap on the first two components while separating on later ones. Formal cluster analysis, run on the PCA scores or on the original data, is the appropriate next step if group assignment matters.

Finally, people sometimes assume that because PCA is “objective” math, the results are interpretation-free. They are not. Choices about which variables to include, whether to standardize, how many components to retain, and whether to rotate all shape the output. Two analysts studying the same phenomenon can produce different-looking PCA plots through defensible but different preprocessing decisions. Reporting those decisions transparently is what separates a trustworthy PCA from a misleading one.

When You Might Want Something Other Than PCA

PCA is a workhorse, but it is not always the right tool. If your goal is classification rather than exploration, supervised methods that use class labels will generally outperform unsupervised PCA. If the data have a known nonlinear structure, manifold-learning methods like UMAP or t-SNE will often reveal groupings that PCA misses. If the variables are counts or proportions rather than continuous measurements, correspondence analysis or other specialized ordination methods may be more appropriate.

In time-series data, temporal autocorrelation can inflate the variance of slow-changing features, causing PC1 to pick up long-term trends rather than the cross-sectional patterns you may care about. Detrending or using time-aware variants of PCA helps in those settings. The broader lesson is that PCA is a starting point for understanding high-dimensional data, not the final word. Its real power lies in how quickly and transparently it reveals the dominant structure, giving you a foundation for more targeted analyses downstream.