How to Use KEGG: Pathways, Genes, and Database Tools

KEGG (Kyoto Encyclopedia of Genes and Genomes) is a database system that connects genome sequences to biological function, organizing genes, chemical compounds, diseases, and drugs into manually drawn pathway maps and hierarchies. It serves as a reference knowledge base linking the genomic space (KEGG GENES) and the chemical space (KEGG LIGAND) through wiring diagrams of interaction and reaction networks (KEGG PATHWAY).1PubMed Central. From genomics to chemical genomics: new developments in KEGG For anyone working with genomic, transcriptomic, or metabolomic data, KEGG is one of the most widely used resources for making sense of what a pile of gene IDs or protein sequences actually means in a living system. Getting comfortable with its main tools and databases turns a steep learning curve into a practical workflow.

The Core Idea Behind KEGG Pathway Maps

At its heart, KEGG is built around manually curated pathway maps. Unlike databases that rely on automated predictions or machine learning to organize biological information, KEGG pathway maps are drawn by human curators who read published literature and distill it into visual diagrams of how molecules interact inside cells.2PubMed Central. KEGG mapping tools for uncovering hidden features in biological data Each map represents a specific biological process, whether that is glycolysis, a signaling cascade, or a disease mechanism. The boxes on these maps represent gene products (usually proteins), and the lines between them represent reactions, activations, inhibitions, or other relationships.

The practical power of KEGG comes from “pathway mapping,” which is the process of taking your own data and projecting it onto these reference maps. If you have a list of genes from a sequenced genome, or a set of differentially expressed genes from an RNA-seq experiment, you can map them onto KEGG pathways to see which biological systems are active, altered, or missing.3PubMed Central. KEGG for linking genomes to life and the environment Instead of staring at a spreadsheet of thousands of gene names, you get a picture of what those genes are doing together.

Understanding KEGG Orthology and K Numbers

Before you can map anything onto a KEGG pathway, your genes need to be assigned to functional groups. This is where the KEGG Orthology (KO) system comes in. KO links DNA and protein sequences to biological functions and pathways, providing a consistent annotation framework across all domains of life.4Briefings in Bioinformatics. DeepKOALA: a scalable deep learning framework for KEGG Orthology assignment Each functional group gets a “K number” (like K00001 for alcohol dehydrogenase). When your gene gets assigned a K number, KEGG knows which pathway boxes that gene belongs in.

The KO system is what makes KEGG work across species. A gene in a mouse and its equivalent in a human might have different sequence IDs, but if they perform the same biochemical function, they share the same K number. This is the foundation for generating organism-specific pathways: once every gene in a genome has been assigned K numbers, KEGG computationally generates pathway maps and functional hierarchies specific to that organism. Expression data, such as from microarray or RNA-seq experiments, can then be mapped onto those organism-specific pathways to explore what the cell or organism is doing under particular conditions.5Nucleic Acids Research. KEGG for linking genomes to life and the environment

Getting K Numbers Assigned with BlastKOALA and GhostKOALA

If you are working with a newly sequenced genome or a metagenome, your sequences will not already have K numbers. You need an annotation tool to assign them. KEGG provides two automatic servers for this: BlastKOALA and GhostKOALA. Both accept amino acid sequences as input, perform KO assignments to characterize individual gene functions, and reconstruct KEGG pathways, BRITE hierarchies, and KEGG modules to give you a high-level functional picture of the organism or ecosystem.6PubMed. BlastKOALA and GhostKOALA: KEGG Tools for Functional Characterization of Genome and Metagenome Sequences

The practical difference between the two is speed and sensitivity. BlastKOALA uses a more thorough sequence search against a non-redundant KEGG GENES dataset, which makes it slower but more precise. GhostKOALA uses a compressed database and faster algorithm, making it better suited for large metagenome datasets where you have millions of sequences and need results in a reasonable time frame. For a single bacterial genome, BlastKOALA is usually the better choice. For a soil or gut microbiome metagenome with enormous sequence files, GhostKOALA will finish while BlastKOALA is still running.

Using KEGG Mapper to Visualize Your Data

Once you have K numbers or KEGG gene IDs, the next step is visualization. KEGG Mapper is a suite of tools that lets you paint your data onto reference pathway maps. You can search for a specific pathway, enter your list of K numbers or gene IDs, and see which boxes on the map light up. The updated version of KEGG Mapper works alongside improved pathway map and BRITE hierarchy viewers, making it straightforward to explore how your data fits into established biological models.2PubMed Central. KEGG mapping tools for uncovering hidden features in biological data

For richer datasets, third-party visualization tools extend what KEGG Mapper can do. One web-based resource accepts transcriptome, proteome, metabolome, or combinations of these data types and overlays them onto KEGG pathways, letting you see all measured entities in the context of cellular processes.7In Silico Biology: Journal of Biological Systems Modeling and Multi-Scale Simulation. KEGG-Based Pathway Visualization Tool for Complex Omics Data Another tool, PaintOmics, takes complete transcriptomics and metabolomics datasets along with lists of significant changes and paints them directly onto KEGG pathway maps.8Bioinformatics. Paintomics: a web based tool for the joint visualization of transcriptomics and metabolomics data These tools are especially useful when you have multi-omics experiments and want to see gene expression changes alongside metabolite concentration shifts in the same pathway diagram.

The Chemical Side of KEGG

Genes are only half the picture. KEGG also catalogs the small molecules that participate in biological reactions. The LIGAND database is a composite system with three main parts: COMPOUND, which stores information about metabolites and other chemical compounds; REACTION, which collects the substrate-to-product relationships representing metabolic and other reactions; and ENZYME, which holds information about the enzyme molecules that catalyze those reactions.9Nucleic Acids Research. LIGAND: database of chemical compounds and reactions in biological pathways

If you are studying metabolism, drug interactions, or metabolomics data, this is where you will spend time. Each compound has its own KEGG ID (a “C number” like C00031 for D-glucose), and each reaction has an “R number.” These IDs link directly into the pathway maps, so when you look at a metabolic pathway, you can click any compound node to see its chemical structure, its role in other pathways, and which enzymes act on it. This cross-referencing is one of KEGG’s most useful features for metabolomics researchers trying to interpret which pathways are perturbed in their experiments.

Enrichment Analysis and a Common Pitfall

One of the most frequent uses of KEGG is pathway enrichment analysis: you take your list of differentially expressed genes, test whether any KEGG pathways have more of those genes than expected by chance, and report the enriched pathways. Many tools outside KEGG itself (such as DAVID, clusterProfiler, or g:Profiler) use KEGG pathway annotations for this purpose.

A subtle but important methodological point: lumping all your differentially expressed genes together can reduce your ability to find disease-relevant pathways. Genes that function together in a pathway tend to have correlated expression levels, which means a given pathway often has an imbalance between up-regulated and down-regulated genes. When you throw both up and down genes into a single enrichment test, the opposing directions can cancel each other out, diluting the signal. Analyzing up-regulated and down-regulated genes separately is more powerful for identifying pathways truly connected to the biological difference you are studying.10PubMed Central. Separate enrichment analysis of pathways for up- and downregulated genes This finding has been demonstrated across multiple tumor types and holds for both microarray and RNA-seq data. If you are running KEGG enrichment and getting disappointingly few hits, splitting your gene list by direction of change is worth trying before concluding that nothing is going on.

A related consideration involves pathway hierarchy. Standard overrepresentation analysis treats each KEGG pathway as independent, but many pathways share genes and have correlated activity. More advanced approaches like path analysis models attempt to separate the direct effect of a KEGG pathway from indirect effects driven by correlation with related pathways, while also estimating the direction of impact.11Molecular Omics. KEGG-PATH: Kyoto encyclopedia of genes and genomes-based pathway analysis using a path analysis model For time-course experiments or multi-treatment designs, these methods can offer a more nuanced picture than simple enrichment.

Disease, Drug, and Virus Databases

KEGG has expanded well beyond core metabolism and signaling. Its DISEASE and DRUG databases have been improved through systematic analysis of drug labels, creating tighter integration between diseases, drugs, and the underlying molecular networks.12PubMed Central. KEGG: new perspectives on genomes, pathways, diseases and drugs This means you can look up a disease and see which pathways are implicated, or look up a drug and see which pathway nodes it targets. For pharmacology and translational research, this linkage is valuable because it places drugs in the same visual framework as the pathways they modulate.

Particularly interesting is the integration of virus-host interactions. KEGG uses “network variation maps,” which are aligned sets of related network diagrams showing how different viruses manipulate specific cellular signaling pathways. For instance, viruses are known to either inhibit or activate apoptosis (programmed cell death) depending on the virus type, and KEGG captures these differences in variation maps and corresponding pathway maps.13PubMed Central. KEGG: integrating viruses and cellular organisms If you are studying how a particular pathogen hijacks host cellular machinery, these maps let you compare strategies across virus families in a standardized format.

Programmatic Access and Scripting Workflows

For anyone running analyses at scale, clicking through web pages is not going to cut it. KEGG offers a REST API that allows programmatic access to its databases, and several community tools wrap this API for easier use. One Python package, kegg_pull, provides both an application programming interface for Python scripts and a command line interface for shell scripting and data analysis pipelines.14PubMed Central. kegg_pull: a software package for the RESTful access and pulling from the Kyoto Encyclopedia of Gene and Genomes This is useful when you need to pull pathway data for hundreds of organisms, download all compound entries matching certain criteria, or integrate KEGG queries into an automated bioinformatics pipeline.

In the R ecosystem, packages like KEGGREST and pathview serve similar purposes, letting you query KEGG directly from R scripts and overlay expression data onto pathway maps without leaving your analysis environment. The API returns data in a structured text format that is relatively easy to parse, though be aware that KEGG imposes rate limits on automated queries. If you are downloading large volumes of data, batch your requests and build in pauses to avoid being temporarily blocked.

Another practical concern involves file formats. KEGG stores its pathway information in KGML, a proprietary XML format. If you need to import KEGG pathways into other software for network analysis or simulation, you will need a converter. KEGGtranslator is a standalone application that can visualize KGML files and convert them into multiple output formats suitable for other tools and algorithms.15PubMed Central. KEGGtranslator: visualizing and converting the KEGG PATHWAY database to various formats This is particularly relevant for researchers doing quantitative network modeling or using graph-analysis software that expects standard formats like SBML or BioPAX.

Metagenomics Applications

KEGG has become a go-to resource for functional profiling of microbial communities. When you sequence a soil sample, a gut microbiome, or ocean water, you get millions of short DNA fragments from hundreds or thousands of species mixed together. Assigning these fragments to KEGG Orthologs lets you ask not just “who is there” but “what can they do.” Because the roughly 25,000 KOs represent a much smaller space than the total number of individual genes across all species, they provide an ideal grouping level for community-level functional analysis.16PubMed Central. Metagenomic functional profiling: to sketch or not to sketch?

Tools like HUMAnN, DIAMOND, and MG-RAST all use KEGG annotations as part of their functional profiling pipelines. If you are comparing the functional potential of microbiomes across conditions (say, healthy versus diseased gut), mapping to KEGG pathways gives you a standardized vocabulary to describe differences. A finding like “nitrogen fixation genes are enriched in soil samples from treatment A” is meaningful because the KEGG pathway for nitrogen metabolism provides a well-curated reference for what those genes actually do together.

How KEGG Compares to Gene Ontology

Researchers often wonder whether to use KEGG or Gene Ontology (GO) for functional annotation and enrichment. The two systems are complementary rather than competing. KEGG’s strength lies in pathway-level context: it tells you not just what a gene does in isolation but where it fits in a chain of reactions or interactions. GO, by contrast, has a much larger vocabulary of functional terms and a more granular hierarchy, but it describes functions as individual annotations rather than placing them on a map.

An early comparison found that for certain biological systems, particularly in prokaryotes, KO-based annotation provided better coverage than GO in both the number of proteins annotated and the quality of annotations, likely because KEGG historically had a stronger focus on prokaryotic species. The biggest limitation of KO noted at the time was a smaller total number of functional terms compared to GO, though this gap has been narrowing as KO continues to grow.17Bioinformatics. Automated genome annotation and pathway identification using the KEGG Orthology (KO) as a controlled vocabulary

A broader assessment found that KEGG tends to have a relatively small number of pathways per organism, but those pathways are relatively large. When excluding very large, unspecific pathways (those covering more than 250 genes), KEGG pathways annotate only about a third of the human protein-coding genome. Other systems like Reactome and GO Cellular Component also each describe less than half of all genes under the same filter.18PubMed Central. Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments No single annotation system covers all human genes in a functionally informative way, which is a good reason to use KEGG and GO together rather than choosing one exclusively. Running enrichment against both databases often catches pathways that the other misses.

Licensing and Access Considerations

One wrinkle that trips up new users: KEGG is free for academic use through its website, but commercial use and bulk downloads require a subscription license. This licensing model is sometimes a source of frustration in the bioinformatics community, especially for researchers at institutions that have not purchased a license or for those building open-source tools that need to distribute KEGG data. Some third-party tools work around this by querying the KEGG API in real time rather than bundling KEGG data, but this approach depends on the API remaining accessible and is subject to rate limits.

For practical purposes, if you are a graduate student or postdoc at a university, you can freely browse the KEGG website, use BlastKOALA and GhostKOALA, run KEGG Mapper, and query the REST API for your research. If you are building a commercial product or need to redistribute KEGG data, check the current licensing terms at kegg.jp. The Reactome database, which is fully open-source, is sometimes used as a complement or alternative when licensing is a concern, though its pathway coverage and organism range differ from KEGG’s.