BioInformatix Resources

Practice Bioinformatics With Real Biological Data

Explore curated public datasets for RNA-Seq, genomics, variant calling, single-cell analysis, metagenomics, epigenomics, machine learning and structural bioinformatics.

Curated from GEO, SRA, ENCODE, 10x Genomics, NIST GIAB, MGnify and PDB

Start Here

Start With These Datasets

These public biological datasets offer focused entry points for building practical skills without beginning with an overly broad study.

GEO: GSE132415

Arabidopsis Heat Stress RNA-Seq

RNA-SeqBeginner

Six samples and three clearly defined conditions make this a focused introduction to count matrices, differential expression, PCA and heatmaps.

Open Dataset
10x Genomics PBMC 3k

10x Genomics PBMC 3k

Single-Cell RNA-SeqBeginner / Intermediate

A focused dataset for learning single-cell QC, normalization, dimensionality reduction, clustering and marker-gene exploration.

Open Dataset
SRA: SRR37806302

Escherichia coli WGS for Genome Assembly

Genome AssemblyBeginner / Intermediate

Paired-end bacterial whole-genome sequencing data provides a practical route through FASTQ QC, contig generation and assembly statistics.

Open Dataset
PDB: 1CRN

Crambin Protein Structure

Structural BioinformaticsBeginner

A 46-residue, one-chain protein structure makes it approachable for learning PDB organization, residues and molecular visualization.

Open Dataset
GEO: GSE24177

Arabidopsis Drought Stress Expression Dataset

Plant MicroarrayBeginner / Intermediate

Twelve drought and control samples support focused practice in expression analysis, clustering and biological interpretation.

Open Dataset

Skill Router

What Do You Want to Practice?

Choose a skill to jump directly to a suitable bioinformatics practice dataset.

Bioinformatics Data Lab

Find a Dataset for Your Next Analysis

Filter this curated collection of bioinformatics datasets by analysis area, difficulty or original public source.

Category
Difficulty
Source

Showing 12 datasets

RNA-Seq / Plant TranscriptomicsBeginner

Arabidopsis Heat Stress RNA-Seq

GEO: GSE132415

A six-sample plant transcriptomics study with control, heat and recovery conditions.

Organism
Arabidopsis thaliana
Design
Control, Heat and Recovery
Samples
6 total
Replicates
2 per condition
Repository
GEO

Practice

  • RNA-Seq
  • Differential expression
  • Count matrices
  • Heatmaps
  • PCA
  • Interpretation
RNA-Seq / Time-CourseIntermediate

Human Sorbitol Stress RNA-Seq Time Course

GEO: GSE310049

A human HEK293 RNA-Seq time course following sorbitol treatment.

Organism
Homo sapiens
Cell system
HEK293
Samples
14 RNA-Seq samples
Replicates
2 per RNA-Seq time point
Treatment
Sorbitol
Repository
GEO

Practice

  • Time-course RNA-Seq
  • PCA
  • Differential expression
  • Expression dynamics
  • DESeq2-oriented practice
Technical details

Time points and conditions: Control, 1 hour, 3 hours, 6 hours, 9 hours, 12 hours and 24 hours.

Raw FASTQ and processed gene-level data are available through the original repository.

RNA-SeqIntermediate

Human Bronchial Epithelial RNA-Seq

GEO: GSE159489 · BioProject: PRJNA669047

A human bronchial epithelial RNA-Seq study suitable for working with biological replicates and expression data.

Organism
Homo sapiens
Samples
18
Replicates
3 biological replicates per treatment group
Repository
GEO and SRA

Practice

  • FASTQ workflows
  • Expression analysis
  • Biological replicates
  • Differential expression
Data availability
Raw sequencing data are available through SRA, with processed expression data through GEO.
Single-Cell RNA-SeqBeginner / Intermediate

10x Genomics PBMC 3k

10x Genomics PBMC 3k

Peripheral blood mononuclear cells from a healthy human donor for single-cell analysis practice.

Organism
Homo sapiens
Sample
Healthy donor peripheral blood mononuclear cells
Cell count
Approximately 2,700 detected cells
Repository
10x Genomics

Practice

  • Seurat
  • Scanpy
  • Single-cell QC
  • Normalization
  • Dimensionality reduction
  • Clustering
  • Marker genes
  • Cell-type exploration
Microarray / Machine LearningIntermediate

Breast Cancer Microarray Dataset

GEO: GSE2034

A lymph-node-negative breast cancer cohort with outcome information and estrogen receptor status.

Organism
Homo sapiens
Samples
286 breast cancer samples
Metadata
Outcome information and estrogen receptor status
Repository
GEO

Practice

  • Microarray analysis
  • Biomarker discovery
  • Clustering
  • Classification
  • Machine learning
  • Outcome analysis
Variant Calling / BenchmarkingIntermediate / Advanced

Genome in a Bottle HG001 / NA12878

GIAB HG001 / NA12878

A reference benchmark genome with high-confidence benchmark variant calls and high-confidence regions.

Organism
Homo sapiens
Type
Reference benchmark genome
Repository
NIST Genome in a Bottle

Practice

  • Variant calling
  • VCF analysis
  • Benchmarking
  • Precision and recall
  • Truth sets
  • Workflow validation
Genome Assembly / NGSBeginner / Intermediate

Escherichia coli WGS for Genome Assembly

SRA: SRR37806302

Paired-end Illumina whole-genome sequencing data for bacterial genome assembly practice.

Organism
Escherichia coli
Data type
Paired-end Illumina whole-genome sequencing
SRA spots
Approximately 1.38 million
Download size
Approximately 206 MB
Repository
SRA

Practice

  • FASTQ QC
  • Bacterial assembly
  • SPAdes-oriented workflows
  • Contig generation
  • Assembly statistics
ATAC-Seq / EpigenomicsIntermediate

K562 ATAC-Seq Dataset

ENCODE: ENCSR017LGQ

An ATAC-Seq experiment using DMSO-treated K562 cells and two isogenic replicates.

Organism
Homo sapiens
Cell line
K562
Condition
DMSO-treated K562
Design
2 isogenic replicates
Repository
ENCODE

Practice

  • ATAC-Seq QC
  • Alignment
  • Chromatin accessibility
  • Peak calling concepts
  • Regulatory genomics
ChIP-Seq / EpigenomicsIntermediate

K562 H3K27ac ChIP-Seq

ENCODE: ENCSR000AKP

A K562 ChIP-Seq experiment targeting the H3K27ac histone modification.

Organism
Homo sapiens
Cell line
K562
Target
H3K27ac
Repository
ENCODE

Practice

  • ChIP-Seq workflows
  • Regulatory regions
  • Peak analysis
  • Histone modifications
  • Enhancer-associated signal
Metagenomics / MicrobiomeIntermediate / Advanced

Human Gut Microbiome After Antibiotic Exposure

MGnify: MGYS00001175 · ENA: PRJEB8094

Human gut microbiome samples collected around antibiotic treatment.

Samples
72
Design
Before antibiotics, end of treatment and approximately 3 months after treatment
Repository
MGnify

Practice

  • Shotgun metagenomics
  • Community analysis
  • Taxonomic profiling
  • Functional profiling
  • Microbiome changes
  • AMR exploration
Plant Microarray / Stress BiologyBeginner / Intermediate

Arabidopsis Drought Stress Expression Dataset

GEO: GSE24177 · BioProject: PRJNA130099

An Arabidopsis microarray expression study comparing drought and well-watered control conditions.

Organism
Arabidopsis thaliana
Samples
12
Design
Drought and control experiments
Replicates
Biological replicates are present
Repository
GEO

Practice

  • Microarray analysis
  • Plant stress expression
  • Differential analysis
  • Clustering
  • Biological interpretation
Structural BioinformaticsBeginner

Crambin Protein Structure

PDB: 1CRN

A compact protein structure for learning the foundations of structural bioinformatics.

Protein
Crambin
Residues
46 amino-acid residues
Chains
One chain
Repository
Protein Data Bank

Practice

  • PDB structure
  • Protein residues
  • Secondary structure
  • Molecular visualization
  • Structural basics

No datasets match the current filters. Try clearing a filter or using a broader search term.

Curated Practice Sets

Featured Collections

Use these collections to build related skills across several public datasets.

RNA-Seq Practice Collection

Practice experimental design, expression dynamics, biological replicates, PCA and differential expression.

Expression & Machine Learning Collection

Practice expression analysis, clustering, biomarker discovery, classification and biological interpretation.

Specialized Data Collection

Move into single-cell analysis, microbiome profiling and protein structure exploration.

Plan Your Workflow

Choose the Right Starting Point

The appropriate starting point depends on whether you want to practice a complete sequencing workflow or focus on downstream analysis.

Raw Data

Best when you want to practice the full workflow, including data handling, quality control and upstream processing.

FASTQBAMSRA

Processed Data

Best when you want to focus on downstream analysis, visualization, statistics and biological interpretation.

Count matrixExpression matrixVCFMetadata tables

From Data to Skills

How to Use the Datasets

Turn public biological datasets into a documented bioinformatics training project.

  1. Choose a dataset
  2. Read the original repository metadata
  3. Download the appropriate files
  4. Follow the relevant tutorial or Learning Path
  5. Perform the analysis
  6. Document your workflow
  7. Interpret the results

Keep your work reproducible. Save the materials another learner would need to understand and repeat your analysis.

  • Scripts
  • README files
  • Figures
  • Environment information
  • Analysis notes

Related BioInformatix Resources

Learn Before You Analyze

Review a concept, follow a structured Learning Path or find an analysis tool before starting your dataset.

Free Tutorials

Review practical bioinformatics concepts and workflows.

Explore Tutorials

Beginner Path

Build the foundations needed to begin working with biological data.

View Beginner Path

Bioinformatics Analyst Path

Develop broader analysis and interpretation skills.

View Analyst Path

NGS Analyst Path

Prepare for sequencing, assembly, variant and epigenomics workflows.

View NGS Path

RNA-Seq Analyst Path

Follow a structured route through RNA-Seq analysis.

View RNA-Seq Path

AI + Bioinformatics Path

Connect biological datasets with machine learning practice.

View AI Path

Bioinformatics Tools

Find tools for analysis, visualization and biological interpretation.

Explore Tools

Courses

Explore guided learning options when you need more structure.

Explore Courses

Public Dataset Usage

The datasets listed here are hosted by their original public repositories and data providers. BioInformatix curates these resources for educational and research practice.

Users should review the original repository metadata, licensing terms, consent conditions and access requirements before using a dataset.

For human genomic data, some repositories may impose additional usage or controlled-access requirements.

Choose a Dataset and Start Practicing

Real progress in bioinformatics comes from working with real biological data. Choose a dataset, follow a workflow and document what you learn.