Protein Data Bank Tutorial: How to Search, View and Download Protein Structures (2026)

  • Home
  • / Protein Data Bank Tutorial: How to Search, View and Download Protein Structures (2026)
Protein Data Bank tutorial

This Protein Data Bank tutorial will show you how to search for experimentally determined three-dimensional structures of proteins, DNA, RNA and biological complexes, understand PDB identifiers, evaluate structural information and download structures for further bioinformatics analysis.

The Protein Data Bank, commonly known as the PDB, is one of the most important resources in structural biology and protein bioinformatics. Researchers use PDB structures to investigate protein function, identify active sites, examine mutations, analyze protein-ligand interactions, perform molecular docking, compare protein structures and validate computational models.

For beginners, however, a PDB record can initially appear complicated. A single structure may contain several protein chains, ligands, metal ions, engineered mutations, missing residues, experimental information and multiple downloadable file formats.

This beginner-friendly Protein Data Bank tutorial explains how to use the PDB correctly and how it connects with other major biological databases including UniProt, Ensembl, NCBI and GenBank.

The global structural archive is coordinated by the Worldwide Protein Data Bank (wwPDB), while the RCSB Protein Data Bank provides one of the principal interfaces for searching, visualizing and analyzing PDB structures.

If you are new to protein databases, first read our UniProt Tutorial: How to Search, Read and Download Protein Data.


What Is the Protein Data Bank? A Beginner’s Protein Data Bank Tutorial

The Protein Data Bank (PDB) is the global archive for experimentally determined three-dimensional structures of biological macromolecules.

PDB structures can include:

  • Proteins
  • DNA
  • RNA
  • Protein-protein complexes
  • Protein-DNA complexes
  • Protein-RNA complexes
  • Antibodies
  • Enzymes
  • Membrane proteins
  • Ribosomes
  • Viral proteins
  • Protein-ligand complexes
  • Large macromolecular assemblies

Unlike UniProt, which primarily provides protein sequence and functional information, the Protein Data Bank focuses on three-dimensional molecular structure.

Structures deposited in PDB are primarily generated using experimental structural biology techniques such as:

  • X-ray crystallography
  • Cryogenic electron microscopy (cryo-EM)
  • Nuclear magnetic resonance spectroscopy (NMR)

One of the main goals of this Protein Data Bank tutorial is to help beginners understand that PDB is not simply a collection of protein images. Each entry represents a structural study containing molecular coordinates, experimental metadata, biological assemblies, sequence information and validation information.


Why Is the Protein Data Bank Important?

Knowing the amino acid sequence of a protein is valuable, but protein function is strongly influenced by its three-dimensional structure.

A protein structure can help researchers investigate:

  • Where an active site is located
  • How substrates bind to enzymes
  • How drugs interact with proteins
  • Which residues form a binding pocket
  • Whether a mutation could disrupt structural stability
  • How two proteins interact
  • Which structural domains make up a protein
  • How a protein changes conformation
  • How conserved residues relate to protein structure

Because of these applications, PDB is widely used in:

  • Structural bioinformatics
  • Molecular biology
  • Drug discovery
  • Molecular docking
  • Protein engineering
  • Enzyme research
  • Genetics
  • Computational biology
  • Pharmacology
  • Cancer research

For a broader introduction to biological databases and computational biology, read What Is Bioinformatics? A Complete Beginner’s Guide.


PDB vs UniProt: What Is the Difference?

Beginners often confuse PDB and UniProt because both contain protein-related information.

However, they have different primary purposes.

FeatureUniProtProtein Data Bank
Protein sequenceYesYes, for deposited structures
Protein functionExtensiveLimited/integrated
3D structureLinks to structuresPrimary focus
Protein domainsYesStructural representation
Post-translational modificationsYesMay be represented
Experimental structureLinks to PDBYes
VariantsYesMay appear in individual structures
LigandsLimitedDetailed structural information
Atomic coordinatesNoYes
Structural visualizationIntegrated/linkedYes

A simplified workflow might look like:

Gene → Ensembl/NCBI → Protein → UniProt → Experimental Structure → PDB

You can learn how to retrieve and interpret protein information in our complete UniProt tutorial.

Protein sequence and functional annotations can also be verified through the official UniProt resource.


What Is a PDB ID?

Every deposited PDB structure receives a unique identifier.

Traditional PDB identifiers contain four characters, for example:

1CRN

or:

4HHB

A PDB identifier refers to a specific experimentally determined structural entry, not necessarily to the complete biological protein.

This distinction is extremely important.

A single protein may have many PDB entries because researchers may determine structures:

  • Using different experimental methods
  • With different ligands
  • With different mutations
  • At different resolutions
  • In different conformational states
  • With different interacting proteins
  • Using different protein fragments or domains

Therefore, finding a PDB entry is only the beginning. You still need to determine whether that structure is appropriate for your research question.


Why Can One Protein Have Many PDB Structures?

Imagine researchers are studying a receptor protein.

One structure might represent:

Receptor alone

Another:

Receptor + inhibitor

Another:

Receptor + activating ligand

Another:

Mutant receptor

And another:

Receptor + interacting protein

Each structure may answer a different biological question.

For this reason, you should never automatically choose the first search result.

The best structure depends on what you are trying to investigate.


Protein Data Bank Tutorial: How to Search for Protein Structures

In this section of our Protein Data Bank tutorial, we will walk through the basic process of finding an appropriate experimental structure rather than simply selecting the first result returned by a search.

A common starting point is the official RCSB Protein Data Bank.

Suppose you want to investigate human TP53.

Step 1: Search by Protein or Gene Name

You could search:

TP53 human

or:

tumor protein p53 Homo sapiens

You can also search using:

  • Protein name
  • Gene symbol
  • UniProt accession
  • PDB identifier
  • Organism
  • Ligand name
  • Experimental method
  • Protein sequence

Using a UniProt accession is particularly useful because it connects a well-defined protein record with corresponding structural entries.


Step 2: Apply Filters

A popular protein may have dozens or hundreds of available structures.

Useful filters include:

  • Organism
  • Experimental method
  • Resolution
  • Polymer type
  • Protein name
  • Ligand
  • Release date
  • Sequence identity
  • Enzyme classification

Filtering helps remove irrelevant structures and makes it easier to identify the structure that best matches your experiment.


Step 3: Examine the Search Results

Before selecting a structure, inspect:

  • PDB ID
  • Structure title
  • Experimental method
  • Resolution
  • Organism
  • Protein chains
  • Ligands
  • Mutations
  • Release date

Do not choose a structure based only on its title.


How to Read a PDB Entry

An important part of any Protein Data Bank tutorial is learning how to evaluate a structure after finding it. A PDB entry contains considerably more information than the three-dimensional model shown in the viewer.

Important sections include the following.


Structure Title

The title describes the experiment or structure that was determined.

A title may indicate:

  • An enzyme bound to an inhibitor
  • A receptor-ligand complex
  • A mutant protein
  • An antibody-antigen complex
  • A protein domain
  • A macromolecular complex

Always read the title before downloading the structure.


Experimental Method

The experimental method tells you how the structure was determined.

X-Ray Crystallography

X-ray crystallography has historically been one of the major techniques for obtaining atomic structures of proteins and other macromolecules.

Researchers crystallize the molecule, collect X-ray diffraction data and use these data to reconstruct its three-dimensional structure.

Cryo-Electron Microscopy

Cryogenic electron microscopy, or cryo-EM, has become particularly important for studying:

  • Large proteins
  • Protein complexes
  • Membrane proteins
  • Ribosomes
  • Viral particles

Cryo-EM has made it possible to study many macromolecular systems that are difficult to crystallize.

NMR Spectroscopy

Nuclear magnetic resonance spectroscopy is commonly used for smaller proteins, peptides and molecular interactions.

NMR entries may contain an ensemble of conformations rather than a single coordinate model.


What Does Resolution Mean in PDB?

For structural techniques where resolution is reported, it provides an important indication of the level of structural detail.

Resolution is normally expressed in angstroms:

Å

One angstrom corresponds to approximately:

10⁻¹⁰ meters

As a broad principle, lower numerical resolution values indicate greater structural detail.

ResolutionGeneral interpretation
Around 1 ÅVery high structural detail
Around 2 ÅHigh-quality atomic detail
Around 2–3 ÅFrequently useful structural detail
Around 3–4 ÅLower structural detail
Above 4 ÅInterpretation generally requires more caution

However, resolution alone does not tell you whether a structure is appropriate for your project.

You must also consider:

  • Sequence coverage
  • Missing residues
  • Ligand state
  • Experimental method
  • Mutations
  • Biological assembly
  • Validation information
  • Your specific research objective

Understanding Protein Chains

A PDB structure can contain multiple chains.

For example:

Chain A

Chain B

Chain C

Different chains might represent:

  • Several copies of the same protein
  • Different proteins
  • Antibody heavy and light chains
  • Receptor and ligand
  • Protein and peptide
  • Components of a larger molecular complex

Before beginning structural analysis, determine which chain corresponds to the protein you want to study.

This is particularly important for:

  • Molecular docking
  • Mutation mapping
  • Structural alignment
  • Protein-protein interaction analysis
  • Active-site analysis

What Is a Biological Assembly?

Another important concept is the difference between the deposited coordinate representation and the biological assembly.

A protein may function as:

  • Monomer
  • Dimer
  • Trimer
  • Tetramer
  • Larger oligomeric complex

For example, a protein may function biologically as a homodimer even if the initially displayed structural representation contains a different arrangement.

If your research involves molecular interactions or protein function, inspect the biological assembly carefully.


Understanding Ligands in PDB

PDB structures commonly contain molecules other than the main protein.

These may include:

  • Drugs
  • Substrates
  • Cofactors
  • Metal ions
  • Nucleotides
  • Sugars
  • Inhibitors
  • Buffer molecules
  • Crystallization components
  • Water molecules

Ligands are particularly important for structural bioinformatics and drug discovery.

For example, a structure containing:

Protein + experimentally bound inhibitor

may be more informative for a docking study than an apo structure containing only the protein.


Active Sites and Binding Sites

Protein structures can reveal the spatial organization of biologically important residues.

Researchers frequently examine:

  • Catalytic residues
  • Ligand-binding residues
  • Metal-binding sites
  • Protein-protein interfaces
  • DNA-binding residues
  • RNA-binding regions

UniProt can help identify experimentally annotated functional residues, while PDB allows you to examine those residues in three-dimensional space.

This Protein Data Bank tutorial therefore also demonstrates why PDB and UniProt should frequently be used together.


Missing Residues in PDB Structures

A common beginner mistake is assuming that a PDB structure represents the complete protein.

Many experimental structures contain only part of the full biological sequence.

Reasons include:

  • Protein flexibility
  • Intrinsic disorder
  • Experimental limitations
  • Construct design
  • Protein degradation
  • Deliberate isolation of one domain

For example, a protein may contain 700 amino acids while a particular PDB structure contains only residues 250–500.

Always compare:

Full biological protein sequence

with:

Sequence represented in the selected PDB structure

Our UniProt tutorial explains how to retrieve the full protein sequence.


Engineered Mutations

Researchers often modify proteins before structural determination.

A PDB structure may contain:

  • Amino acid substitutions
  • Deletions
  • Truncations
  • Stabilizing mutations
  • Affinity tags
  • Fusion proteins
  • Engineered constructs

This is especially important if you intend to use the structure for:

  • Variant interpretation
  • Molecular docking
  • Protein modeling
  • Functional analysis

Never assume that every PDB structure represents the natural wild-type protein.


How to View Protein Structures Online

The RCSB PDB interface provides interactive three-dimensional structural visualization.

You can use the structure viewer to:

  • Rotate the molecule
  • Zoom in and out
  • Select chains
  • Highlight amino acids
  • Display ligands
  • Examine secondary structure
  • Investigate molecular surfaces
  • Measure distances
  • Inspect neighboring residues

Common visualization styles include:

Cartoon Representation

Useful for visualizing:

  • Alpha helices
  • Beta sheets
  • Overall protein fold

Surface Representation

Useful for:

  • Binding pockets
  • Protein interfaces
  • Molecular shape

Stick Representation

Useful for:

  • Side chains
  • Ligands
  • Active-site residues

Sphere Representation

Useful for emphasizing particular atoms or molecules.


Protein Secondary Structure

Proteins commonly contain secondary structural elements such as:

  • Alpha helices
  • Beta sheets
  • Loops
  • Turns

Structural visualization helps researchers understand how these elements combine to create a functional protein fold.

Secondary structure can also help explain:

  • Protein stability
  • Domain organization
  • Binding sites
  • Mutation effects

Protein Data Bank Tutorial: How to Download PDB Structures

The next step in this Protein Data Bank tutorial is downloading the correct structural file. The format you choose depends on the downstream software and type of structural analysis you plan to perform.

After identifying an appropriate PDB entry, you can download structural information in formats such as:

  • PDB
  • PDBx/mmCIF
  • BinaryCIF
  • Experimental-data formats

What Is the PDB File Format?

Traditional PDB files contain atomic coordinates and structural information in a text-based format.

A simplified coordinate record may resemble:

ATOM 1 N MET A 1 20.154 34.245 18.271

The information can include:

  • Atom number
  • Atom name
  • Residue name
  • Chain
  • Residue number
  • X coordinate
  • Y coordinate
  • Z coordinate
  • Occupancy
  • Temperature factor

Traditional PDB files remain compatible with many structural bioinformatics and molecular visualization applications.


What Is PDBx/mmCIF?

The PDBx/mmCIF format is the modern archival format used for PDB structural data.

It supports richer and more complex structural information than the traditional fixed-column PDB format.

mmCIF is particularly useful when working with:

  • Large protein complexes
  • Ribosomes
  • Viral particles
  • Large macromolecular assemblies
  • Modern structural bioinformatics pipelines

PDB vs mmCIF: Which Format Should You Use?

For small proteins and older software, traditional PDB files remain convenient.

For modern analyses, mmCIF is frequently the better option because it avoids several limitations of the older PDB format.

The correct choice depends on the software you plan to use.


Can You Download Protein Sequences from PDB?

Yes, PDB entries contain sequence information for the macromolecules represented in the structure.

However, if your primary objective is retrieving the complete biological protein sequence, UniProt is generally preferable.

Why?

A PDB protein sequence may represent:

  • Only one domain
  • A truncated protein
  • An engineered mutant
  • A crystallization construct
  • A fusion protein
  • Only part of the full biological sequence

Use PDB when you want the experimentally studied structural construct.

Use UniProt when you want a comprehensive protein sequence record.


Protein Data Bank vs AlphaFold Database

This comparison has become increasingly important in structural bioinformatics.

Protein Data Bank

PDB primarily archives experimentally determined macromolecular structures.

These may be determined using:

  • X-ray crystallography
  • Cryo-EM
  • NMR spectroscopy

AlphaFold Database

The AlphaFold Protein Structure Database provides computationally predicted protein structures.

Predicted structures can be extremely useful when an appropriate experimental structure is unavailable.

However, computationally predicted structures should not automatically be treated as equivalent to experimentally determined structures.

Before selecting a predicted structure, ask:

Is a suitable experimental structure available?

If yes, evaluate the experimental structure first.

If no suitable experimental structure exists, a high-confidence predicted model may provide an alternative depending on the objective of the study.

A future BioInformatix article in this cluster should cover:

AlphaFold Tutorial: How to Search, Interpret and Download Predicted Protein Structures


Protein Data Bank Tutorial: How to Choose the Best Structure

When many structures are available for the same protein, structure selection requires more than simply choosing the highest-resolution entry.

A useful workflow is:

  1. Confirm the correct organism.
  2. Confirm the correct protein or UniProt accession.
  3. Examine sequence coverage.
  4. Check for engineered mutations.
  5. Examine the experimental method.
  6. Evaluate resolution where applicable.
  7. Check whether the relevant ligand is present.
  8. Identify the correct protein chain.
  9. Examine missing residues.
  10. Check the biological assembly.
  11. Review structural validation information.
  12. Read the associated publication.

There is no universal “best PDB structure.”

The best structure is the one that is most appropriate for your research question.


Choosing a PDB Structure for Molecular Docking

For learners interested in molecular docking, this Protein Data Bank tutorial provides the structure-selection foundation. Receptor preparation, protonation, ligand preparation, docking configuration and validation require additional steps.

For docking applications, evaluate:

  • Relevant protein conformation
  • Experimental resolution
  • Presence of a co-crystallized ligand
  • Completeness of the binding site
  • Missing residues
  • Engineered mutations
  • Cofactors
  • Metal ions
  • Water molecules
  • Protein chains

A protein structure containing an experimentally bound ligand can help identify a biologically relevant binding site.

Before docking, researchers may also need to:

  • Remove irrelevant molecules
  • Add hydrogens
  • Correct protonation states
  • Repair missing atoms
  • Select the correct chains
  • Retain important cofactors
  • Prepare the receptor

Do not automatically delete every heteroatom from a structure. Some metals, cofactors or structural molecules may be biologically essential.


Using PDB for Mutation Analysis

Suppose genomic analysis identifies a missense variant such as:

p.Arg175His

A protein structure can help determine whether the affected residue occurs within:

  • A catalytic region
  • Protein core
  • DNA-binding domain
  • Protein-protein interface
  • Ligand-binding pocket
  • Conserved structural domain
  • Flexible loop

Structural analysis can therefore provide additional biological context for variants.

However, structural location alone is not sufficient to classify a genetic variant as pathogenic or benign.

For genomic and transcript-level variant analysis, see our Ensembl Genome Browser Guide.


Using PDB for Protein-Protein Interaction Analysis

Structures containing multiple proteins can reveal:

  • Molecular interfaces
  • Hydrogen bonds
  • Salt bridges
  • Hydrophobic contacts
  • Interface residues
  • Binding orientations

This is particularly valuable when studying:

  • Signaling proteins
  • Receptors
  • Antibodies
  • Transcription factors
  • Enzyme complexes

Combining PDB with Protein Sequence Analysis

PDB and UniProt can be combined into a powerful sequence-structure workflow:

Retrieve protein sequence from UniProt

Identify domains and functional residues

Search PDB for experimental structures

Compare sequence coverage

Select the appropriate structure

Map conserved or functional residues

Map mutations

Perform structural analysis

You can learn practical protein sequence and structural analysis in our Protein Bioinformatics Masterclass: Sequence Analysis, Structure Prediction, Modeling & Proteomics.


Searching PDB Using a Protein Sequence

Sometimes you have a protein sequence but do not know whether an experimental structure exists.

Sequence similarity searching can help identify PDB structures related to your sequence.

A general workflow is:

Protein sequence

Sequence similarity search

Matching PDB structures

Evaluate sequence identity

Evaluate sequence coverage

Inspect organism and structure

Select suitable structure

Good sequence identity and coverage may indicate that an existing structure could be useful.

However, you should still evaluate:

  • Organism
  • Protein domain coverage
  • Mutations
  • Ligands
  • Structure completeness
  • Biological state

Downloading Multiple PDB Structures

Large-scale structural studies may require hundreds or thousands of structures.

Instead of manually downloading every entry, researchers can use:

  • APIs
  • Python
  • Linux shell scripts
  • Bulk-download resources
  • Bioinformatics workflows

Applications include:

  • Structural bioinformatics
  • Machine learning
  • Protein-family analysis
  • Comparative structural studies
  • Ligand screening

To develop the programming skills required for automated biological-data retrieval, explore Learn Bioinformatics Data Analysis: Master Python, Linux and R Scripting.

You may also find these tutorials helpful:


Connecting PDB with Other Biological Databases

The Protein Data Bank becomes much more powerful when integrated with other biological databases.

PDB + UniProt

Use UniProt for:

  • Protein sequence
  • Protein function
  • Domains
  • Active sites
  • Variants
  • Isoforms

Use PDB for:

  • Experimental structure
  • Atomic coordinates
  • Bound ligands
  • Structural complexes

Read our UniProt Tutorial.

PDB + Ensembl

Use Ensembl to connect the protein with:

  • Gene
  • Transcript
  • Exons
  • Genomic coordinates
  • Genetic variants

Read our Ensembl Genome Browser Guide.

PDB + NCBI

NCBI provides:

  • Gene information
  • Nucleotide sequences
  • Scientific literature
  • Genome assemblies
  • Genetic variation resources

Read What Is NCBI? A Complete Beginner’s Guide.

PDB + GenBank

GenBank provides nucleotide sequence records associated with genes, transcripts and genomes.

Read our GenBank Complete Guide.


Common Protein Data Bank Mistakes Beginners Should Avoid

Before finishing this Protein Data Bank tutorial, it is important to understand several mistakes that frequently cause problems in structural bioinformatics projects.

Choosing the First Search Result

Several structures may exist for the same protein. Compare them carefully.

Ignoring the Organism

Proteins from different organisms may be structurally similar but biologically different.

Ignoring Protein Coverage

A PDB structure may represent only one domain of a much larger protein.

Ignoring Engineered Mutations

The deposited construct may contain substitutions or deletions that are absent from the natural protein.

Selecting a Structure Only by Resolution

Resolution matters, but structural completeness and biological relevance may matter more for a specific experiment.

Ignoring Ligands

A ligand-bound structure and an apo structure can represent different protein conformations.

Ignoring Missing Residues

Missing residues near a binding pocket or protein interface can significantly affect downstream analysis.

Using the Wrong Chain

A PDB structure may contain several different proteins.

Confusing Experimental and Predicted Structures

PDB is centered on experimental structural data, whereas resources such as AlphaFoldDB provide computational predictions.

Removing Every Heteroatom

Metal ions, cofactors and other molecules may be required for biological activity.

Ignoring the Biological Assembly

The biologically relevant complex may differ from the initially displayed structural coordinates.


Additional Protein Data Bank Learning Resources

For additional structural biology learning materials, explore PDB-101, an educational resource maintained by RCSB PDB.

It provides accessible information about:

  • Protein structures
  • Biological molecules
  • Molecular visualization
  • Structural biology
  • Important macromolecular systems

These external scientific resources complement practical bioinformatics training and help beginners understand the biological context behind PDB structures.


Continue Learning After This Protein Data Bank Tutorial

Reading about PDB provides the foundation, but practical experience is required to develop structural bioinformatics skills.

Start with our free:

Introduction to Biological Databases for Bioinformatics

This course introduces important biological data resources and teaches beginners how to retrieve and interpret biological information.

For dedicated protein analysis, continue with:

Protein Bioinformatics Masterclass: Sequence Analysis, Structure Prediction, Modeling & Proteomics

The course covers practical concepts related to:

  • Protein sequence analysis
  • Physicochemical properties
  • Protein domains and motifs
  • Protein structure prediction
  • Homology modeling
  • Structural validation
  • Protein visualization
  • Proteomics
  • Structural bioinformatics

Beginners can also start with our free:

Roadmap to Bioinformatics: A Beginner’s Guide to Careers, Skills & Learning Path

For complete project-based training, explore:

Learn Bioinformatics: Beginner to Master Through Real-World Projects

To learn Python, R and Linux for automated biological-data analysis, continue with:

Learn Bioinformatics Data Analysis: Master Python, Linux and R Scripting


Recommended Learning Path for Protein Bioinformatics

If you are beginning protein bioinformatics, follow this sequence:

  1. Read What Is Bioinformatics? A Complete Beginner’s Guide.
  2. Complete Introduction to Biological Databases for Bioinformatics.
  3. Study our UniProt Tutorial.
  4. Practice downloading complete protein sequences.
  5. Search PDB for corresponding experimental structures.
  6. Compare sequence coverage between UniProt and PDB.
  7. Examine protein chains and biological assemblies.
  8. Identify ligands and functional residues.
  9. Download PDB or mmCIF files.
  10. Practice molecular visualization.
  11. Progress to protein modeling and structural analysis with the Protein Bioinformatics Masterclass.

Frequently Asked Questions

What does PDB stand for?

PDB stands for Protein Data Bank.

Is the Protein Data Bank free?

Yes. PDB structural data can be searched, viewed and downloaded freely.

What is a PDB ID?

A PDB ID uniquely identifies a structural entry deposited in the Protein Data Bank.

Does PDB contain protein sequences?

Yes. PDB entries contain sequence information for the macromolecules represented in a structure. However, UniProt is generally more appropriate when you need a complete biological protein sequence.

Does PDB contain predicted protein structures?

The core PDB archive focuses on experimentally determined macromolecular structures. Computational predictions are available through resources such as AlphaFoldDB.

What is the difference between PDB and UniProt?

UniProt primarily focuses on protein sequences and functional annotations. PDB focuses on experimentally determined three-dimensional molecular structures.

Which PDB structure should I choose?

Evaluate organism, sequence coverage, mutation status, experimental method, resolution, ligand state, missing residues, biological assembly and your specific research objective.

What is better: PDB or mmCIF?

Traditional PDB format remains widely supported, but PDBx/mmCIF is the modern archival format and is better suited to complex and large structures.

Can I use PDB structures for molecular docking?

Yes. PDB structures are frequently used for docking, but the receptor must be carefully selected and prepared before analysis.

What is an apo protein structure?

An apo structure generally represents a protein without its relevant bound ligand or cofactor.

What is a holo structure?

A holo structure generally represents a protein in a ligand-bound or cofactor-bound state.

What does 2 Å resolution mean?

It indicates structural information at approximately a two-angstrom scale. For applicable experimental methods, lower numerical resolution values generally provide greater structural detail.

Can PDB structures contain mutations?

Yes. Experimental structures frequently contain engineered substitutions, truncations or other modifications.

Should I use PDB or AlphaFold?

Use an appropriate high-quality experimental PDB structure when one is available and suitable for your research question. Predicted AlphaFold structures can be valuable when experimental structures are unavailable or incomplete, but prediction confidence and biological context should be evaluated carefully.


Final Thoughts

This Protein Data Bank tutorial provides the foundation needed to search, evaluate, visualize and download experimentally determined macromolecular structures correctly.

The Protein Data Bank allows researchers to move beyond amino acid sequences and examine biological molecules in three dimensions. Learning how to search PDB, understand PDB identifiers, evaluate experimental methods, examine resolution, identify chains, inspect ligands and select appropriate structural files is an essential part of modern structural bioinformatics.

Before selecting any structure, evaluate:

  • Correct protein
  • Correct organism
  • Sequence coverage
  • Experimental method
  • Resolution
  • Engineered mutations
  • Ligands
  • Missing residues
  • Protein chains
  • Biological assembly
  • Structural validation
  • Relevance to your scientific question

By combining the concepts covered in this Protein Data Bank tutorial with UniProt, Ensembl, NCBI and GenBank, you can build a complete workflow connecting genes, transcripts, protein sequences, functions and three-dimensional molecular structures.

These foundations prepare you for more advanced topics including protein modeling, molecular docking, mutation analysis, drug discovery and computational structural biology.


Bioinformatix Team

BioInformatix is an online bioinformatics training platform focused on providing practical education in genomics, transcriptomics, computational biology, artificial intelligence, and biological data analysis. We help students, researchers, and professionals build industry-ready skills through hands-on projects, real-world datasets, and career-focused learning programs.

BIOINFORMATICS

Protein Data Bank Tutorial: How to Search, View and Download Protein Structures (2026)

This Protein Data Bank tutorial will show you how to search for experimentally determined three-dimensional structures of proteins, DNA, RNA and biological complexes, understand PDB identifiers, evaluate structural information and download structures for further bioinformatics analysis. The Protein Data Bank, commonly known as the PDB, is one of the most important resources in structural biology […]

19 min read Reading time
Aug 7, 2026 Published
Protein Data Bank Tutorial: How to Search, View and Download Protein Structures (2026)
BIOINFORMATIX GUIDE Learn the Concept. Apply the Workflow.
ARTICLE CONTENTS On This Page
Reading progress 0%