BioInformatix Resources
Practice Bioinformatics With Real Biological Data
Explore curated public datasets for RNA-Seq, genomics, variant calling, single-cell analysis, metagenomics, epigenomics, machine learning and structural bioinformatics.
Start Here
Start With These Datasets
These public biological datasets offer focused entry points for building practical skills without beginning with an overly broad study.
Arabidopsis Heat Stress RNA-Seq
Six samples and three clearly defined conditions make this a focused introduction to count matrices, differential expression, PCA and heatmaps.
Open Dataset10x Genomics PBMC 3k
A focused dataset for learning single-cell QC, normalization, dimensionality reduction, clustering and marker-gene exploration.
Open DatasetEscherichia coli WGS for Genome Assembly
Paired-end bacterial whole-genome sequencing data provides a practical route through FASTQ QC, contig generation and assembly statistics.
Open DatasetCrambin Protein Structure
A 46-residue, one-chain protein structure makes it approachable for learning PDB organization, residues and molecular visualization.
Open DatasetArabidopsis Drought Stress Expression Dataset
Twelve drought and control samples support focused practice in expression analysis, clustering and biological interpretation.
Open DatasetSkill Router
What Do You Want to Practice?
Choose a skill to jump directly to a suitable bioinformatics practice dataset.
Bioinformatics Data Lab
Find a Dataset for Your Next Analysis
Filter this curated collection of bioinformatics datasets by analysis area, difficulty or original public source.
Arabidopsis Heat Stress RNA-Seq
GEO: GSE132415A six-sample plant transcriptomics study with control, heat and recovery conditions.
Practice
- RNA-Seq
- Differential expression
- Count matrices
- Heatmaps
- PCA
- Interpretation
Human Sorbitol Stress RNA-Seq Time Course
GEO: GSE310049A human HEK293 RNA-Seq time course following sorbitol treatment.
Practice
- Time-course RNA-Seq
- PCA
- Differential expression
- Expression dynamics
- DESeq2-oriented practice
Technical details
Time points and conditions: Control, 1 hour, 3 hours, 6 hours, 9 hours, 12 hours and 24 hours.
Raw FASTQ and processed gene-level data are available through the original repository.
Human Bronchial Epithelial RNA-Seq
GEO: GSE159489 · BioProject: PRJNA669047A human bronchial epithelial RNA-Seq study suitable for working with biological replicates and expression data.
Practice
- FASTQ workflows
- Expression analysis
- Biological replicates
- Differential expression
Data availability
10x Genomics PBMC 3k
10x Genomics PBMC 3kPeripheral blood mononuclear cells from a healthy human donor for single-cell analysis practice.
Practice
- Seurat
- Scanpy
- Single-cell QC
- Normalization
- Dimensionality reduction
- Clustering
- Marker genes
- Cell-type exploration
Breast Cancer Microarray Dataset
GEO: GSE2034A lymph-node-negative breast cancer cohort with outcome information and estrogen receptor status.
Practice
- Microarray analysis
- Biomarker discovery
- Clustering
- Classification
- Machine learning
- Outcome analysis
Genome in a Bottle HG001 / NA12878
GIAB HG001 / NA12878A reference benchmark genome with high-confidence benchmark variant calls and high-confidence regions.
Practice
- Variant calling
- VCF analysis
- Benchmarking
- Precision and recall
- Truth sets
- Workflow validation
Escherichia coli WGS for Genome Assembly
SRA: SRR37806302Paired-end Illumina whole-genome sequencing data for bacterial genome assembly practice.
Practice
- FASTQ QC
- Bacterial assembly
- SPAdes-oriented workflows
- Contig generation
- Assembly statistics
K562 ATAC-Seq Dataset
ENCODE: ENCSR017LGQAn ATAC-Seq experiment using DMSO-treated K562 cells and two isogenic replicates.
Practice
- ATAC-Seq QC
- Alignment
- Chromatin accessibility
- Peak calling concepts
- Regulatory genomics
K562 H3K27ac ChIP-Seq
ENCODE: ENCSR000AKPA K562 ChIP-Seq experiment targeting the H3K27ac histone modification.
Practice
- ChIP-Seq workflows
- Regulatory regions
- Peak analysis
- Histone modifications
- Enhancer-associated signal
Human Gut Microbiome After Antibiotic Exposure
MGnify: MGYS00001175 · ENA: PRJEB8094Human gut microbiome samples collected around antibiotic treatment.
Practice
- Shotgun metagenomics
- Community analysis
- Taxonomic profiling
- Functional profiling
- Microbiome changes
- AMR exploration
Arabidopsis Drought Stress Expression Dataset
GEO: GSE24177 · BioProject: PRJNA130099An Arabidopsis microarray expression study comparing drought and well-watered control conditions.
Practice
- Microarray analysis
- Plant stress expression
- Differential analysis
- Clustering
- Biological interpretation
Crambin Protein Structure
PDB: 1CRNA compact protein structure for learning the foundations of structural bioinformatics.
Practice
- PDB structure
- Protein residues
- Secondary structure
- Molecular visualization
- Structural basics
No datasets match the current filters. Try clearing a filter or using a broader search term.
Curated Practice Sets
Featured Collections
Use these collections to build related skills across several public datasets.
RNA-Seq Practice Collection
Practice experimental design, expression dynamics, biological replicates, PCA and differential expression.
Genomics Practice Collection
Explore variant benchmarking, bacterial assembly, chromatin accessibility and histone modification analysis.
Expression & Machine Learning Collection
Practice expression analysis, clustering, biomarker discovery, classification and biological interpretation.
Specialized Data Collection
Move into single-cell analysis, microbiome profiling and protein structure exploration.
Plan Your Workflow
Choose the Right Starting Point
The appropriate starting point depends on whether you want to practice a complete sequencing workflow or focus on downstream analysis.
Raw Data
Best when you want to practice the full workflow, including data handling, quality control and upstream processing.
Processed Data
Best when you want to focus on downstream analysis, visualization, statistics and biological interpretation.
From Data to Skills
How to Use the Datasets
Turn public biological datasets into a documented bioinformatics training project.
- Choose a dataset
- Read the original repository metadata
- Download the appropriate files
- Follow the relevant tutorial or Learning Path
- Perform the analysis
- Document your workflow
- Interpret the results
Keep your work reproducible. Save the materials another learner would need to understand and repeat your analysis.
- Scripts
- README files
- Figures
- Environment information
- Analysis notes
Related BioInformatix Resources
Learn Before You Analyze
Review a concept, follow a structured Learning Path or find an analysis tool before starting your dataset.
Free Tutorials
Review practical bioinformatics concepts and workflows.
Explore TutorialsBeginner Path
Build the foundations needed to begin working with biological data.
View Beginner PathBioinformatics Analyst Path
Develop broader analysis and interpretation skills.
View Analyst PathNGS Analyst Path
Prepare for sequencing, assembly, variant and epigenomics workflows.
View NGS PathRNA-Seq Analyst Path
Follow a structured route through RNA-Seq analysis.
View RNA-Seq PathAI + Bioinformatics Path
Connect biological datasets with machine learning practice.
View AI PathBioinformatics Tools
Find tools for analysis, visualization and biological interpretation.
Explore ToolsCourses
Explore guided learning options when you need more structure.
Explore CoursesPublic Dataset Usage
The datasets listed here are hosted by their original public repositories and data providers. BioInformatix curates these resources for educational and research practice.
Users should review the original repository metadata, licensing terms, consent conditions and access requirements before using a dataset.
For human genomic data, some repositories may impose additional usage or controlled-access requirements.
Choose a Dataset and Start Practicing
Real progress in bioinformatics comes from working with real biological data. Choose a dataset, follow a workflow and document what you learn.

