
This Protein Data Bank tutorial will show you how to search for experimentally determined three-dimensional structures of proteins, DNA, RNA and biological complexes, understand PDB identifiers, evaluate structural information and download structures for further bioinformatics analysis.
The Protein Data Bank, commonly known as the PDB, is one of the most important resources in structural biology and protein bioinformatics. Researchers use PDB structures to investigate protein function, identify active sites, examine mutations, analyze protein-ligand interactions, perform molecular docking, compare protein structures and validate computational models.
For beginners, however, a PDB record can initially appear complicated. A single structure may contain several protein chains, ligands, metal ions, engineered mutations, missing residues, experimental information and multiple downloadable file formats.
This beginner-friendly Protein Data Bank tutorial explains how to use the PDB correctly and how it connects with other major biological databases including UniProt, Ensembl, NCBI and GenBank.
The global structural archive is coordinated by the Worldwide Protein Data Bank (wwPDB), while the RCSB Protein Data Bank provides one of the principal interfaces for searching, visualizing and analyzing PDB structures.
If you are new to protein databases, first read our UniProt Tutorial: How to Search, Read and Download Protein Data.
What Is the Protein Data Bank? A Beginner’s Protein Data Bank Tutorial
The Protein Data Bank (PDB) is the global archive for experimentally determined three-dimensional structures of biological macromolecules.
PDB structures can include:
- Proteins
- DNA
- RNA
- Protein-protein complexes
- Protein-DNA complexes
- Protein-RNA complexes
- Antibodies
- Enzymes
- Membrane proteins
- Ribosomes
- Viral proteins
- Protein-ligand complexes
- Large macromolecular assemblies
Unlike UniProt, which primarily provides protein sequence and functional information, the Protein Data Bank focuses on three-dimensional molecular structure.
Structures deposited in PDB are primarily generated using experimental structural biology techniques such as:
- X-ray crystallography
- Cryogenic electron microscopy (cryo-EM)
- Nuclear magnetic resonance spectroscopy (NMR)
One of the main goals of this Protein Data Bank tutorial is to help beginners understand that PDB is not simply a collection of protein images. Each entry represents a structural study containing molecular coordinates, experimental metadata, biological assemblies, sequence information and validation information.
Why Is the Protein Data Bank Important?
Knowing the amino acid sequence of a protein is valuable, but protein function is strongly influenced by its three-dimensional structure.
A protein structure can help researchers investigate:
- Where an active site is located
- How substrates bind to enzymes
- How drugs interact with proteins
- Which residues form a binding pocket
- Whether a mutation could disrupt structural stability
- How two proteins interact
- Which structural domains make up a protein
- How a protein changes conformation
- How conserved residues relate to protein structure
Because of these applications, PDB is widely used in:
- Structural bioinformatics
- Molecular biology
- Drug discovery
- Molecular docking
- Protein engineering
- Enzyme research
- Genetics
- Computational biology
- Pharmacology
- Cancer research
For a broader introduction to biological databases and computational biology, read What Is Bioinformatics? A Complete Beginner’s Guide.
PDB vs UniProt: What Is the Difference?
Beginners often confuse PDB and UniProt because both contain protein-related information.
However, they have different primary purposes.
| Feature | UniProt | Protein Data Bank |
|---|---|---|
| Protein sequence | Yes | Yes, for deposited structures |
| Protein function | Extensive | Limited/integrated |
| 3D structure | Links to structures | Primary focus |
| Protein domains | Yes | Structural representation |
| Post-translational modifications | Yes | May be represented |
| Experimental structure | Links to PDB | Yes |
| Variants | Yes | May appear in individual structures |
| Ligands | Limited | Detailed structural information |
| Atomic coordinates | No | Yes |
| Structural visualization | Integrated/linked | Yes |
A simplified workflow might look like:
Gene → Ensembl/NCBI → Protein → UniProt → Experimental Structure → PDB
You can learn how to retrieve and interpret protein information in our complete UniProt tutorial.
Protein sequence and functional annotations can also be verified through the official UniProt resource.
What Is a PDB ID?
Every deposited PDB structure receives a unique identifier.
Traditional PDB identifiers contain four characters, for example:
1CRN
or:
4HHB
A PDB identifier refers to a specific experimentally determined structural entry, not necessarily to the complete biological protein.
This distinction is extremely important.
A single protein may have many PDB entries because researchers may determine structures:
- Using different experimental methods
- With different ligands
- With different mutations
- At different resolutions
- In different conformational states
- With different interacting proteins
- Using different protein fragments or domains
Therefore, finding a PDB entry is only the beginning. You still need to determine whether that structure is appropriate for your research question.
Why Can One Protein Have Many PDB Structures?
Imagine researchers are studying a receptor protein.
One structure might represent:
Receptor alone
Another:
Receptor + inhibitor
Another:
Receptor + activating ligand
Another:
Mutant receptor
And another:
Receptor + interacting protein
Each structure may answer a different biological question.
For this reason, you should never automatically choose the first search result.
The best structure depends on what you are trying to investigate.
Protein Data Bank Tutorial: How to Search for Protein Structures
In this section of our Protein Data Bank tutorial, we will walk through the basic process of finding an appropriate experimental structure rather than simply selecting the first result returned by a search.
A common starting point is the official RCSB Protein Data Bank.
Suppose you want to investigate human TP53.
Step 1: Search by Protein or Gene Name
You could search:
TP53 human
or:
tumor protein p53 Homo sapiens
You can also search using:
- Protein name
- Gene symbol
- UniProt accession
- PDB identifier
- Organism
- Ligand name
- Experimental method
- Protein sequence
Using a UniProt accession is particularly useful because it connects a well-defined protein record with corresponding structural entries.
Step 2: Apply Filters
A popular protein may have dozens or hundreds of available structures.
Useful filters include:
- Organism
- Experimental method
- Resolution
- Polymer type
- Protein name
- Ligand
- Release date
- Sequence identity
- Enzyme classification
Filtering helps remove irrelevant structures and makes it easier to identify the structure that best matches your experiment.
Step 3: Examine the Search Results
Before selecting a structure, inspect:
- PDB ID
- Structure title
- Experimental method
- Resolution
- Organism
- Protein chains
- Ligands
- Mutations
- Release date
Do not choose a structure based only on its title.
How to Read a PDB Entry
An important part of any Protein Data Bank tutorial is learning how to evaluate a structure after finding it. A PDB entry contains considerably more information than the three-dimensional model shown in the viewer.
Important sections include the following.
Structure Title
The title describes the experiment or structure that was determined.
A title may indicate:
- An enzyme bound to an inhibitor
- A receptor-ligand complex
- A mutant protein
- An antibody-antigen complex
- A protein domain
- A macromolecular complex
Always read the title before downloading the structure.
Experimental Method
The experimental method tells you how the structure was determined.
X-Ray Crystallography
X-ray crystallography has historically been one of the major techniques for obtaining atomic structures of proteins and other macromolecules.
Researchers crystallize the molecule, collect X-ray diffraction data and use these data to reconstruct its three-dimensional structure.
Cryo-Electron Microscopy
Cryogenic electron microscopy, or cryo-EM, has become particularly important for studying:
- Large proteins
- Protein complexes
- Membrane proteins
- Ribosomes
- Viral particles
Cryo-EM has made it possible to study many macromolecular systems that are difficult to crystallize.
NMR Spectroscopy
Nuclear magnetic resonance spectroscopy is commonly used for smaller proteins, peptides and molecular interactions.
NMR entries may contain an ensemble of conformations rather than a single coordinate model.
What Does Resolution Mean in PDB?
For structural techniques where resolution is reported, it provides an important indication of the level of structural detail.
Resolution is normally expressed in angstroms:
Å
One angstrom corresponds to approximately:
10⁻¹⁰ meters
As a broad principle, lower numerical resolution values indicate greater structural detail.
| Resolution | General interpretation |
|---|---|
| Around 1 Å | Very high structural detail |
| Around 2 Å | High-quality atomic detail |
| Around 2–3 Å | Frequently useful structural detail |
| Around 3–4 Å | Lower structural detail |
| Above 4 Å | Interpretation generally requires more caution |
However, resolution alone does not tell you whether a structure is appropriate for your project.
You must also consider:
- Sequence coverage
- Missing residues
- Ligand state
- Experimental method
- Mutations
- Biological assembly
- Validation information
- Your specific research objective
Understanding Protein Chains
A PDB structure can contain multiple chains.
For example:
Chain A
Chain B
Chain C
Different chains might represent:
- Several copies of the same protein
- Different proteins
- Antibody heavy and light chains
- Receptor and ligand
- Protein and peptide
- Components of a larger molecular complex
Before beginning structural analysis, determine which chain corresponds to the protein you want to study.
This is particularly important for:
- Molecular docking
- Mutation mapping
- Structural alignment
- Protein-protein interaction analysis
- Active-site analysis
What Is a Biological Assembly?
Another important concept is the difference between the deposited coordinate representation and the biological assembly.
A protein may function as:
- Monomer
- Dimer
- Trimer
- Tetramer
- Larger oligomeric complex
For example, a protein may function biologically as a homodimer even if the initially displayed structural representation contains a different arrangement.
If your research involves molecular interactions or protein function, inspect the biological assembly carefully.
Understanding Ligands in PDB
PDB structures commonly contain molecules other than the main protein.
These may include:
- Drugs
- Substrates
- Cofactors
- Metal ions
- Nucleotides
- Sugars
- Inhibitors
- Buffer molecules
- Crystallization components
- Water molecules
Ligands are particularly important for structural bioinformatics and drug discovery.
For example, a structure containing:
Protein + experimentally bound inhibitor
may be more informative for a docking study than an apo structure containing only the protein.
Active Sites and Binding Sites
Protein structures can reveal the spatial organization of biologically important residues.
Researchers frequently examine:
- Catalytic residues
- Ligand-binding residues
- Metal-binding sites
- Protein-protein interfaces
- DNA-binding residues
- RNA-binding regions
UniProt can help identify experimentally annotated functional residues, while PDB allows you to examine those residues in three-dimensional space.
This Protein Data Bank tutorial therefore also demonstrates why PDB and UniProt should frequently be used together.
Missing Residues in PDB Structures
A common beginner mistake is assuming that a PDB structure represents the complete protein.
Many experimental structures contain only part of the full biological sequence.
Reasons include:
- Protein flexibility
- Intrinsic disorder
- Experimental limitations
- Construct design
- Protein degradation
- Deliberate isolation of one domain
For example, a protein may contain 700 amino acids while a particular PDB structure contains only residues 250–500.
Always compare:
Full biological protein sequence
with:
Sequence represented in the selected PDB structure
Our UniProt tutorial explains how to retrieve the full protein sequence.
Engineered Mutations
Researchers often modify proteins before structural determination.
A PDB structure may contain:
- Amino acid substitutions
- Deletions
- Truncations
- Stabilizing mutations
- Affinity tags
- Fusion proteins
- Engineered constructs
This is especially important if you intend to use the structure for:
- Variant interpretation
- Molecular docking
- Protein modeling
- Functional analysis
Never assume that every PDB structure represents the natural wild-type protein.
How to View Protein Structures Online
The RCSB PDB interface provides interactive three-dimensional structural visualization.
You can use the structure viewer to:
- Rotate the molecule
- Zoom in and out
- Select chains
- Highlight amino acids
- Display ligands
- Examine secondary structure
- Investigate molecular surfaces
- Measure distances
- Inspect neighboring residues
Common visualization styles include:
Cartoon Representation
Useful for visualizing:
- Alpha helices
- Beta sheets
- Overall protein fold
Surface Representation
Useful for:
- Binding pockets
- Protein interfaces
- Molecular shape
Stick Representation
Useful for:
- Side chains
- Ligands
- Active-site residues
Sphere Representation
Useful for emphasizing particular atoms or molecules.
Protein Secondary Structure
Proteins commonly contain secondary structural elements such as:
- Alpha helices
- Beta sheets
- Loops
- Turns
Structural visualization helps researchers understand how these elements combine to create a functional protein fold.
Secondary structure can also help explain:
- Protein stability
- Domain organization
- Binding sites
- Mutation effects
Protein Data Bank Tutorial: How to Download PDB Structures
The next step in this Protein Data Bank tutorial is downloading the correct structural file. The format you choose depends on the downstream software and type of structural analysis you plan to perform.
After identifying an appropriate PDB entry, you can download structural information in formats such as:
- PDB
- PDBx/mmCIF
- BinaryCIF
- Experimental-data formats
What Is the PDB File Format?
Traditional PDB files contain atomic coordinates and structural information in a text-based format.
A simplified coordinate record may resemble:
ATOM 1 N MET A 1 20.154 34.245 18.271
The information can include:
- Atom number
- Atom name
- Residue name
- Chain
- Residue number
- X coordinate
- Y coordinate
- Z coordinate
- Occupancy
- Temperature factor
Traditional PDB files remain compatible with many structural bioinformatics and molecular visualization applications.
What Is PDBx/mmCIF?
The PDBx/mmCIF format is the modern archival format used for PDB structural data.
It supports richer and more complex structural information than the traditional fixed-column PDB format.
mmCIF is particularly useful when working with:
- Large protein complexes
- Ribosomes
- Viral particles
- Large macromolecular assemblies
- Modern structural bioinformatics pipelines
PDB vs mmCIF: Which Format Should You Use?
For small proteins and older software, traditional PDB files remain convenient.
For modern analyses, mmCIF is frequently the better option because it avoids several limitations of the older PDB format.
The correct choice depends on the software you plan to use.
Can You Download Protein Sequences from PDB?
Yes, PDB entries contain sequence information for the macromolecules represented in the structure.
However, if your primary objective is retrieving the complete biological protein sequence, UniProt is generally preferable.
Why?
A PDB protein sequence may represent:
- Only one domain
- A truncated protein
- An engineered mutant
- A crystallization construct
- A fusion protein
- Only part of the full biological sequence
Use PDB when you want the experimentally studied structural construct.
Use UniProt when you want a comprehensive protein sequence record.
Protein Data Bank vs AlphaFold Database
This comparison has become increasingly important in structural bioinformatics.
Protein Data Bank
PDB primarily archives experimentally determined macromolecular structures.
These may be determined using:
- X-ray crystallography
- Cryo-EM
- NMR spectroscopy
AlphaFold Database
The AlphaFold Protein Structure Database provides computationally predicted protein structures.
Predicted structures can be extremely useful when an appropriate experimental structure is unavailable.
However, computationally predicted structures should not automatically be treated as equivalent to experimentally determined structures.
Before selecting a predicted structure, ask:
Is a suitable experimental structure available?
If yes, evaluate the experimental structure first.
If no suitable experimental structure exists, a high-confidence predicted model may provide an alternative depending on the objective of the study.
A future BioInformatix article in this cluster should cover:
AlphaFold Tutorial: How to Search, Interpret and Download Predicted Protein Structures
Protein Data Bank Tutorial: How to Choose the Best Structure
When many structures are available for the same protein, structure selection requires more than simply choosing the highest-resolution entry.
A useful workflow is:
- Confirm the correct organism.
- Confirm the correct protein or UniProt accession.
- Examine sequence coverage.
- Check for engineered mutations.
- Examine the experimental method.
- Evaluate resolution where applicable.
- Check whether the relevant ligand is present.
- Identify the correct protein chain.
- Examine missing residues.
- Check the biological assembly.
- Review structural validation information.
- Read the associated publication.
There is no universal “best PDB structure.”
The best structure is the one that is most appropriate for your research question.
Choosing a PDB Structure for Molecular Docking
For learners interested in molecular docking, this Protein Data Bank tutorial provides the structure-selection foundation. Receptor preparation, protonation, ligand preparation, docking configuration and validation require additional steps.
For docking applications, evaluate:
- Relevant protein conformation
- Experimental resolution
- Presence of a co-crystallized ligand
- Completeness of the binding site
- Missing residues
- Engineered mutations
- Cofactors
- Metal ions
- Water molecules
- Protein chains
A protein structure containing an experimentally bound ligand can help identify a biologically relevant binding site.
Before docking, researchers may also need to:
- Remove irrelevant molecules
- Add hydrogens
- Correct protonation states
- Repair missing atoms
- Select the correct chains
- Retain important cofactors
- Prepare the receptor
Do not automatically delete every heteroatom from a structure. Some metals, cofactors or structural molecules may be biologically essential.
Using PDB for Mutation Analysis
Suppose genomic analysis identifies a missense variant such as:
p.Arg175His
A protein structure can help determine whether the affected residue occurs within:
- A catalytic region
- Protein core
- DNA-binding domain
- Protein-protein interface
- Ligand-binding pocket
- Conserved structural domain
- Flexible loop
Structural analysis can therefore provide additional biological context for variants.
However, structural location alone is not sufficient to classify a genetic variant as pathogenic or benign.
For genomic and transcript-level variant analysis, see our Ensembl Genome Browser Guide.
Using PDB for Protein-Protein Interaction Analysis
Structures containing multiple proteins can reveal:
- Molecular interfaces
- Hydrogen bonds
- Salt bridges
- Hydrophobic contacts
- Interface residues
- Binding orientations
This is particularly valuable when studying:
- Signaling proteins
- Receptors
- Antibodies
- Transcription factors
- Enzyme complexes
Combining PDB with Protein Sequence Analysis
PDB and UniProt can be combined into a powerful sequence-structure workflow:
Retrieve protein sequence from UniProt
↓
Identify domains and functional residues
↓
Search PDB for experimental structures
↓
Compare sequence coverage
↓
Select the appropriate structure
↓
Map conserved or functional residues
↓
Map mutations
↓
Perform structural analysis
You can learn practical protein sequence and structural analysis in our Protein Bioinformatics Masterclass: Sequence Analysis, Structure Prediction, Modeling & Proteomics.
Searching PDB Using a Protein Sequence
Sometimes you have a protein sequence but do not know whether an experimental structure exists.
Sequence similarity searching can help identify PDB structures related to your sequence.
A general workflow is:
Protein sequence
↓
Sequence similarity search
↓
Matching PDB structures
↓
Evaluate sequence identity
↓
Evaluate sequence coverage
↓
Inspect organism and structure
↓
Select suitable structure
Good sequence identity and coverage may indicate that an existing structure could be useful.
However, you should still evaluate:
- Organism
- Protein domain coverage
- Mutations
- Ligands
- Structure completeness
- Biological state
Downloading Multiple PDB Structures
Large-scale structural studies may require hundreds or thousands of structures.
Instead of manually downloading every entry, researchers can use:
- APIs
- Python
- Linux shell scripts
- Bulk-download resources
- Bioinformatics workflows
Applications include:
- Structural bioinformatics
- Machine learning
- Protein-family analysis
- Comparative structural studies
- Ligand screening
To develop the programming skills required for automated biological-data retrieval, explore Learn Bioinformatics Data Analysis: Master Python, Linux and R Scripting.
You may also find these tutorials helpful:
- Python for Bioinformatics: A Complete Beginner’s Guide
- Linux for Bioinformatics: The Complete Beginner’s Guide
- R Programming for Bioinformatics: A Complete Beginner’s Guide
Connecting PDB with Other Biological Databases
The Protein Data Bank becomes much more powerful when integrated with other biological databases.
PDB + UniProt
Use UniProt for:
- Protein sequence
- Protein function
- Domains
- Active sites
- Variants
- Isoforms
Use PDB for:
- Experimental structure
- Atomic coordinates
- Bound ligands
- Structural complexes
Read our UniProt Tutorial.
PDB + Ensembl
Use Ensembl to connect the protein with:
- Gene
- Transcript
- Exons
- Genomic coordinates
- Genetic variants
Read our Ensembl Genome Browser Guide.
PDB + NCBI
NCBI provides:
- Gene information
- Nucleotide sequences
- Scientific literature
- Genome assemblies
- Genetic variation resources
Read What Is NCBI? A Complete Beginner’s Guide.
PDB + GenBank
GenBank provides nucleotide sequence records associated with genes, transcripts and genomes.
Read our GenBank Complete Guide.
Common Protein Data Bank Mistakes Beginners Should Avoid
Before finishing this Protein Data Bank tutorial, it is important to understand several mistakes that frequently cause problems in structural bioinformatics projects.
Choosing the First Search Result
Several structures may exist for the same protein. Compare them carefully.
Ignoring the Organism
Proteins from different organisms may be structurally similar but biologically different.
Ignoring Protein Coverage
A PDB structure may represent only one domain of a much larger protein.
Ignoring Engineered Mutations
The deposited construct may contain substitutions or deletions that are absent from the natural protein.
Selecting a Structure Only by Resolution
Resolution matters, but structural completeness and biological relevance may matter more for a specific experiment.
Ignoring Ligands
A ligand-bound structure and an apo structure can represent different protein conformations.
Ignoring Missing Residues
Missing residues near a binding pocket or protein interface can significantly affect downstream analysis.
Using the Wrong Chain
A PDB structure may contain several different proteins.
Confusing Experimental and Predicted Structures
PDB is centered on experimental structural data, whereas resources such as AlphaFoldDB provide computational predictions.
Removing Every Heteroatom
Metal ions, cofactors and other molecules may be required for biological activity.
Ignoring the Biological Assembly
The biologically relevant complex may differ from the initially displayed structural coordinates.
Additional Protein Data Bank Learning Resources
For additional structural biology learning materials, explore PDB-101, an educational resource maintained by RCSB PDB.
It provides accessible information about:
- Protein structures
- Biological molecules
- Molecular visualization
- Structural biology
- Important macromolecular systems
These external scientific resources complement practical bioinformatics training and help beginners understand the biological context behind PDB structures.
Continue Learning After This Protein Data Bank Tutorial
Reading about PDB provides the foundation, but practical experience is required to develop structural bioinformatics skills.
Start with our free:
Introduction to Biological Databases for Bioinformatics
This course introduces important biological data resources and teaches beginners how to retrieve and interpret biological information.
For dedicated protein analysis, continue with:
Protein Bioinformatics Masterclass: Sequence Analysis, Structure Prediction, Modeling & Proteomics
The course covers practical concepts related to:
- Protein sequence analysis
- Physicochemical properties
- Protein domains and motifs
- Protein structure prediction
- Homology modeling
- Structural validation
- Protein visualization
- Proteomics
- Structural bioinformatics
Beginners can also start with our free:
Roadmap to Bioinformatics: A Beginner’s Guide to Careers, Skills & Learning Path
For complete project-based training, explore:
Learn Bioinformatics: Beginner to Master Through Real-World Projects
To learn Python, R and Linux for automated biological-data analysis, continue with:
Learn Bioinformatics Data Analysis: Master Python, Linux and R Scripting
Recommended Learning Path for Protein Bioinformatics
If you are beginning protein bioinformatics, follow this sequence:
- Read What Is Bioinformatics? A Complete Beginner’s Guide.
- Complete Introduction to Biological Databases for Bioinformatics.
- Study our UniProt Tutorial.
- Practice downloading complete protein sequences.
- Search PDB for corresponding experimental structures.
- Compare sequence coverage between UniProt and PDB.
- Examine protein chains and biological assemblies.
- Identify ligands and functional residues.
- Download PDB or mmCIF files.
- Practice molecular visualization.
- Progress to protein modeling and structural analysis with the Protein Bioinformatics Masterclass.
Frequently Asked Questions
What does PDB stand for?
PDB stands for Protein Data Bank.
Is the Protein Data Bank free?
Yes. PDB structural data can be searched, viewed and downloaded freely.
What is a PDB ID?
A PDB ID uniquely identifies a structural entry deposited in the Protein Data Bank.
Does PDB contain protein sequences?
Yes. PDB entries contain sequence information for the macromolecules represented in a structure. However, UniProt is generally more appropriate when you need a complete biological protein sequence.
Does PDB contain predicted protein structures?
The core PDB archive focuses on experimentally determined macromolecular structures. Computational predictions are available through resources such as AlphaFoldDB.
What is the difference between PDB and UniProt?
UniProt primarily focuses on protein sequences and functional annotations. PDB focuses on experimentally determined three-dimensional molecular structures.
Which PDB structure should I choose?
Evaluate organism, sequence coverage, mutation status, experimental method, resolution, ligand state, missing residues, biological assembly and your specific research objective.
What is better: PDB or mmCIF?
Traditional PDB format remains widely supported, but PDBx/mmCIF is the modern archival format and is better suited to complex and large structures.
Can I use PDB structures for molecular docking?
Yes. PDB structures are frequently used for docking, but the receptor must be carefully selected and prepared before analysis.
What is an apo protein structure?
An apo structure generally represents a protein without its relevant bound ligand or cofactor.
What is a holo structure?
A holo structure generally represents a protein in a ligand-bound or cofactor-bound state.
What does 2 Å resolution mean?
It indicates structural information at approximately a two-angstrom scale. For applicable experimental methods, lower numerical resolution values generally provide greater structural detail.
Can PDB structures contain mutations?
Yes. Experimental structures frequently contain engineered substitutions, truncations or other modifications.
Should I use PDB or AlphaFold?
Use an appropriate high-quality experimental PDB structure when one is available and suitable for your research question. Predicted AlphaFold structures can be valuable when experimental structures are unavailable or incomplete, but prediction confidence and biological context should be evaluated carefully.
Final Thoughts
This Protein Data Bank tutorial provides the foundation needed to search, evaluate, visualize and download experimentally determined macromolecular structures correctly.
The Protein Data Bank allows researchers to move beyond amino acid sequences and examine biological molecules in three dimensions. Learning how to search PDB, understand PDB identifiers, evaluate experimental methods, examine resolution, identify chains, inspect ligands and select appropriate structural files is an essential part of modern structural bioinformatics.
Before selecting any structure, evaluate:
- Correct protein
- Correct organism
- Sequence coverage
- Experimental method
- Resolution
- Engineered mutations
- Ligands
- Missing residues
- Protein chains
- Biological assembly
- Structural validation
- Relevance to your scientific question
By combining the concepts covered in this Protein Data Bank tutorial with UniProt, Ensembl, NCBI and GenBank, you can build a complete workflow connecting genes, transcripts, protein sequences, functions and three-dimensional molecular structures.
These foundations prepare you for more advanced topics including protein modeling, molecular docking, mutation analysis, drug discovery and computational structural biology.


