Graph neural networks (GNNs) are a family of deep learning models built to work directly on graph-structured data, where information lives not just in individual items but in the connections between them. Think social networks, molecules, supply chains, or the circuits inside a computer chip. Traditional neural networks expect data arranged in neat grids or sequences, but GNNs take in a web of nodes and edges and learn from the pattern of relationships itself. That ability to reason over structure makes them one of the most consequential developments in machine learning over the past decade, with deployments ranging from drug discovery to billion-scale recommendation engines.
Why Graphs Need Their Own Neural Networks
Most data people encounter every day fits neatly into rows and columns, or into ordered sequences like text. Images, for example, are grids of pixels, and standard convolutional neural networks (CNNs) were designed specifically for that grid layout. But a huge amount of real-world information is relational: who follows whom, which atoms bond to which, which roads connect which cities. Forcing that kind of data into a flat table means throwing away the very structure that makes it meaningful.
Standard convolutional networks struggle here because they rely on fixed, regularly spaced filters that assume a grid-like geometry. When data lives in irregular, non-grid arrangements, those filters cannot capture the spatial relationships between elements properly.1Expert Systems with Applications. Non-Euclidean Spectral-Spatial feature mining network with Gated GCN-CNN for hyperspectral image classification Graphs are inherently “non-Euclidean,” meaning the usual notion of distance on a flat plane does not apply. Embedding graph data into standard flat representations can distort the true distances between nodes, especially in graphs whose connectivity follows a power-law distribution, where a few nodes have vastly more connections than most.2PubMed. Hyperbolic Bernstein Neural Networks: Enhancing graph convolutions in non-Euclidean spaces GNNs solve this by operating on the graph’s own topology, respecting the connections as they actually exist rather than hammering them into a shape that suits older architectures.
How Message Passing Works
At their core, most GNNs follow a framework called message passing. The idea is intuitive: every node in a graph collects information from its neighbors, combines that information with what it already knows about itself, and updates its own internal representation. Then the process repeats. After several rounds, each node’s representation reflects not just its own features but the broader neighborhood it sits in. A GNN takes in a graph and outputs a graph, progressively transforming the information stored at nodes, edges, and the graph level without changing the underlying connectivity.3Distill. A Gentle Introduction to Graph Neural Networks – Section: Graph Neural Networks
Imagine a molecule represented as a graph, where atoms are nodes and chemical bonds are edges. In the first round of message passing, each atom “hears” only from the atoms directly bonded to it. After a second round, it has indirect information from atoms two bonds away. Stack enough rounds, and every atom has a sense of the molecule’s overall shape. The model then uses these enriched representations to make predictions, like whether the molecule will bind to a particular protein target.
Different GNN architectures vary in how they aggregate neighbor information. Some treat all neighbors equally, some learn to weight certain neighbors more heavily (an attention mechanism), and some operate in the frequency domain of the graph’s structure rather than the spatial domain. Research has shown that these spectral and spatial approaches are more closely related than they first appear, and can be analyzed within a unified framework.4arXiv. Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks
Where GNNs Are Already Deployed at Scale
The clearest proof that GNNs matter is that major technology companies already run them in production on enormous datasets. Pinterest developed PinSage, a graph convolutional network trained on a graph with roughly 3 billion nodes and 18 billion edges representing pins, boards, and the relationships between them. The system generates personalized recommendations by learning embeddings that capture both the content of a pin and its position in the broader graph of user-curated boards. In offline tests, user studies, and A/B experiments, PinSage produced higher-quality recommendations than alternative deep learning and graph-based approaches.5arXiv. Graph Convolutional Neural Networks for Web-Scale Recommender Systems
Pinterest later extended this approach with MultiBiSage, a model that goes beyond the single pin-to-board graph and incorporates multiple types of entities and interactions, such as users, creators, and behaviors like add-to-cart or long-click. Modeling these diverse interactions through multiple bipartite graphs improved embedding quality further.6arXiv. MultiBiSage: A Web-Scale Recommendation System Using Multiple Bipartite Graphs at Pinterest Other companies running GNN-based recommendation systems include social media platforms and e-commerce sites, though many keep implementation details proprietary.
Drug Discovery and Molecular Graphs
Molecules are natural graphs: atoms are nodes, bonds are edges. This makes GNNs a compelling fit for predicting molecular properties, a task central to drug discovery. Rather than relying on hand-crafted chemical descriptors, GNNs can learn representations directly from molecular structure. Multiple studies have found that GNN-based models yield more promising results than traditional descriptor-based methods for molecular property prediction.7PubMed Central. Could graph neural networks learn better molecular representation for drug discovery? A comparison study of descriptor-based and graph-based models
One practical challenge in this space is data scarcity. High-fidelity experimental measurements of molecular properties are expensive and slow to generate, while lower-quality computational estimates are cheap and plentiful. Transfer learning approaches using GNNs can leverage large pools of low-fidelity data to improve predictions on the small, expensive high-fidelity datasets that actually matter for drug development decisions.8PubMed Central. Transfer learning with graph neural networks for improved molecular property prediction in the multi-fidelity setting This is a good example of GNNs not just performing well in a benchmark but addressing a genuine bottleneck in an industry workflow.
Chip Floorplanning
A less obvious but increasingly important application is in semiconductor design. Chip floorplanning, the process of deciding where to physically place circuit components on a chip, is a combinatorial optimization problem that traditionally requires weeks of engineering effort. GNNs can learn the relationship between a circuit’s connectivity graph and the resulting physical performance metrics like wire length and timing.
GraphPlanner, a variational graph convolutional network for floorplanning, was shown to improve placement runtime by about 25% while reducing wire length by roughly 4% on average compared to state-of-the-art mixed-size placers alone.9ACM Transactions on Design Automation of Electronic Systems. GraphPlanner: Floorplanning with Graph Neural Network Google’s work went further, framing chip floorplanning as a reinforcement learning problem and using an edge-based graph convolutional architecture to learn transferable representations of chips. The key insight is that the model gets better and faster at solving new chip designs by drawing on experience from previous ones, a capability no human designer can match at that scale.10Nature. A graph placement methodology for fast chip design
What GNNs Cannot Distinguish
GNNs have a well-understood theoretical ceiling. Their ability to tell two different graphs apart is bounded by a classical algorithm called the Weisfeiler-Lehman (WL) graph isomorphism test. The standard message-passing GNN is exactly as powerful as the basic (1-WL) version of this test: any pair of graphs the WL test can distinguish, a sufficiently well-designed GNN can also distinguish, and vice versa.11PubMed. Weisfeiler-Lehman goes dynamic: An analysis of the expressive power of Graph Neural Networks for attributed and dynamic graphs 12NeurIPS Proceedings. Graph Neural Networks: What They Are & Why They Matter
This matters because the 1-WL test is known to fail on certain graph structures. There exist pairs of non-identical graphs that it (and therefore standard GNNs) will treat as the same. This has spurred research into more powerful GNN variants that go beyond 1-WL expressiveness, higher-order message passing schemes, random feature augmentations, and architectures that incorporate subgraph information. The WL equivalence has become the standard yardstick for measuring and comparing the expressive power of new GNN designs.13arXiv. Expressiveness and Approximation Properties of Graph Neural Networks
Over-Squashing and Over-Smoothing
Two practical problems plague GNNs that stack many message-passing layers. Over-smoothing happens when repeated rounds of neighbor aggregation cause all node representations to converge toward the same values, washing out the distinctions between nodes. Over-squashing is the complementary problem: information from distant parts of the graph gets compressed through bottleneck nodes, losing signal along the way. Both issues stem from the interaction between the message-passing process and the graph’s topology, and they limit how deep GNNs can usefully go.14arXiv. Rewiring Techniques to Mitigate Oversquashing and Oversmoothing in GNNs: A Survey
A growing body of work addresses these problems through graph rewiring, modifying the graph’s structure before or during training to improve information flow. Techniques include adding shortcut edges between distant nodes, removing redundant edges, and using spectral properties of the graph to guide topology changes.15arXiv. Graph Rewiring in GNNs to Mitigate Over-Squashing and Over-Smoothing: A Survey These methods represent an important shift in thinking: rather than only improving the neural network, you can also improve the graph it operates on.
Scaling to Massive Graphs
Web-scale graphs with billions of nodes pose serious computational challenges. A naive GNN that tries to aggregate information from every neighbor at every layer quickly runs out of memory and time. The GraphSAGE approach tackled this by learning a function that generates embeddings through sampling and aggregating features from a node’s local neighborhood, rather than training individual embeddings for each node in the graph.16arXiv. Inductive Representation Learning on Large Graphs This inductive approach means the model can generalize to nodes it has never seen during training, which is essential for dynamic graphs where new nodes appear constantly (new users signing up, new products being listed).
Mini-batch training, neighborhood sampling, and distributed computing frameworks have collectively made it feasible to train GNNs on graphs with billions of edges, as Pinterest’s production systems demonstrate. But scaling remains an active area of engineering. The irregular memory access patterns of graph computation are a poor fit for GPUs optimized around dense, regular tensor operations, and specialized graph-processing hardware is still in early stages.
Heterogeneous and Temporal Graphs
Real-world graphs rarely consist of a single type of node and a single type of edge. A knowledge graph might have person nodes, organization nodes, and location nodes connected by “works at,” “born in,” and “founded” edges. These are heterogeneous graphs, and standard GNNs that treat all nodes and edges identically leave significant information on the table.
The problem gets harder when the graph also changes over time. Heterogeneous temporal graph neural networks tackle both challenges simultaneously, using hierarchical aggregation mechanisms that handle different relationship types separately before combining them, and that also model how relationships evolve across time steps.17arXiv. Heterogeneous Temporal Graph Neural Network More recent work has focused on making these models more efficient, since the combination of heterogeneity and temporal dynamics can multiply the computational cost.18arXiv. Simple and Efficient Heterogeneous Temporal Graph Neural Network
An emerging approach uses mathematical structures called sheaves to model local heterogeneity in how information should flow between different parts of a graph. Rather than applying the same transformation everywhere, sheaf-based GNNs learn different transformations for different local neighborhoods, adapting over time to capture complex spatiotemporal patterns.19arXiv. Dynamic Sheaf Diffusion Networks with Adaptive Local Structure for Heterogeneous Spatio-Temporal Graph Learning
Making GNN Predictions Interpretable
A persistent concern with GNNs, as with deep learning generally, is that their predictions can be opaque. If a GNN flags a molecule as toxic or recommends a particular product, a user or regulator might reasonably ask why. GNNExplainer was the first general, model-agnostic tool for this problem. Given a prediction, it identifies a compact subgraph and a small subset of node features that were most influential in the GNN’s decision.20arXiv. GNNExplainer: Generating Explanations for Graph Neural Networks
Later approaches like PGExplainer improved on this by using a learned neural network to parameterize the explanation process. This means explanations can be generated collectively for groups of instances rather than one at a time, offering better generalization and the ability to work in inductive settings where new, unseen graphs appear at test time. PGExplainer showed up to about 25% relative improvement in explanation quality over earlier baselines on graph classification tasks.21Advances in Neural Information Processing Systems. Parameterized Explainer for Graph Neural Network Interpretability remains an active and practically important area, particularly for applications in healthcare and finance where decisions need to be auditable.
Generating New Graphs
GNNs are not limited to analyzing existing graphs. Generative models built on graph neural network architectures can learn to produce entirely new graphs that resemble a training distribution. This has obvious appeal in chemistry, where generating novel molecular structures with desired properties could accelerate the search for new drugs or materials.
Early work demonstrated that GNN-based generative models can capture both the structure and attributes of graphs, and once trained, produce good quality samples of both synthetic and real molecular graphs, either unconditionally or guided by specific conditions.22arXiv. Learning Deep Generative Models of Graphs MolGAN took a different approach, using an implicit generative model that directly produces molecular graphs without requiring expensive graph matching procedures or assumptions about how nodes should be ordered.23arXiv. MolGAN: An implicit generative model for small molecular graphs The ability to optimize a differentiable model that outputs molecular graphs lets researchers side-step the brute-force search through the vast space of possible chemical structures.
Graph Transformers
The transformer architecture that dominates natural language processing and computer vision has been adapted for graphs, producing a family of models called graph transformers. These models use attention mechanisms to weigh the influence of different nodes, but they face a design tension: a standard transformer treats its input as a fully connected set (every element attends to every other element), which throws away the graph’s sparse connectivity structure. Various strategies have been developed to preserve structural information, including specialized positional encodings, structure-aware attention patterns, and hybrid architectures that combine message passing with transformer layers.24ACM Computing Surveys. A Survey of Graph Transformers: Architectures, Theories and Applications
Some researchers have found that instead of treating the molecule as a fully connected graph (where every atom attends to every other atom), incorporating the actual graph connectivity as an inductive bias through message diffusion mechanisms produces better molecular representations and avoids what might be called a message enrichment explosion.25arXiv. Learning Attributed Graph Representations with Communicative Message Passing Transformer The question of when graph transformers outperform classical GNNs and when they are unnecessary overhead is still being worked out, but the trend is clearly toward combining the strengths of both paradigms.
Adversarial Vulnerability
GNNs inherit the adversarial vulnerability problem from deep learning more broadly, but with a twist. In image classification, an attacker might add imperceptible noise to pixel values. In a graph, an attacker can also manipulate the topology itself: adding or removing edges, injecting fake nodes, or subtly altering node features. Because GNNs propagate information through the graph structure, a small number of strategic edge additions can cascade through the message-passing process and degrade predictions across the graph.26Computers, Materials and Continua. Adversarial Attack Defense in Graph Neural Networks via Multiview Learning and Attention-Guided Topology Filtering
This is not merely an academic concern. In fraud detection, a malicious actor could try to add spurious connections to legitimate accounts to make a fraudulent account appear trustworthy. In recommendation systems, fake interactions could manipulate what gets surfaced to millions of users. Defense strategies include topology filtering (detecting and removing suspicious edges), training with adversarial examples, and multiview learning that cross-checks information from different perspectives of the graph. The arms race between attack and defense methods is ongoing and has real stakes anywhere GNNs touch safety-critical or financially sensitive systems.
Benchmarks and the State of Evaluation
A persistent challenge in the GNN field has been inconsistent evaluation. Different papers used different datasets, different data splits, and different metrics, making it hard to tell whether a new architecture was genuinely better or just evaluated differently. The Open Graph Benchmark (OGB) was created to address this, providing a diverse set of large-scale, realistic benchmark datasets spanning social networks, biological networks, molecular graphs, source code, and knowledge graphs. Each dataset comes with a standardized evaluation protocol, including application-specific data splits that test out-of-distribution generalization rather than random shuffling.27arXiv. Open Graph Benchmark: Datasets for Machine Learning on Graphs
OGB quickly became the standard proving ground, and its leaderboards have been a useful forcing function. Researchers have applied neural architecture search to automatically discover GNN architectures tailored to specific OGB tasks, demonstrating that the best architecture often depends heavily on the dataset and task at hand.28arXiv.org. Graph Property Prediction on Open Graph Benchmark: A Winning Solution by Graph Neural Architecture Search The benchmark experiments also revealed that scalability to large graphs and generalization under realistic data splits remain genuine unsolved problems, not just minor engineering details. For anyone evaluating GNN methods for a practical application, these benchmarks offer the closest thing the field has to an apples-to-apples comparison.