Analysis code for the ATAC-seq and RNA-seq context data characterizing the cell types and
stimulation states used to interpret regulatory variant effects in the accompanying
Alzheimer's-disease MPRA study (ad_mpra_chen, see below). This covers chromatin accessibility
and gene expression profiling of THP1 monocytes/macrophages under cytokine stimulation, primary
human macrophages, microglia (HMC3, iPSC-derived, and public single-cell/ex vivo references),
iPSC-derived neurons, and public comparison cell lines (K562, HepG2, SK-N-SH) — used to annotate
MPRA-tested enhancers with cell-type/state-specific open chromatin and to place THP1/HMC3
stimulation states in the context of human microglial transcriptional states.
This repository holds the analysis notebooks/scripts and sample sheets; the peak calls,
differential-analysis result tables, and count matrices they produce are not tracked here — see
DATA.md for where to get them.
This repository is a companion to the MPRA analysis repository. Several notebooks there load
tables produced by the pipelines in this repository, e.g.
notebooks/07_atac_overlap/7.0.atac_annotation_ml_cluster_overlap.ipynb and
notebooks/misc/differential_enhancer_ATAC_comparison.ipynb (overlap MPRA allele effects with
THP1 differential ATAC peaks). The sequence-based CNN modeling code that consumes this
accessibility data lives in ad_mpra_chen/machinelearning_prepost_processing/, alongside the
MPRA analysis it directly supports.
atac/
differential_analysis/
featurecount_DESEQ2_*.ipynb featureCounts -> DESeq2/IHW differential peak analysis,
one notebook per comparison/cell-type group
download_encode_peaks.sh, rename.sh Fetch/rename public reference peak sets (ENCODE)
sample_sheet*.csv Sample/condition metadata for the notebooks above
jaccard_index/
jaccard_index.ipynb Peak-overlap (Jaccard index) across reference datasets
rna-seq/
THP1/, THP1_mock_electroporation/, THP1_SEC63/, ipsc_exc_neuron/, Kang2019/
featurecount_DESEQ2_*.ipynb featureCounts -> DESeq2/IHW differential expression,
fgsea, per dataset
CRISPRi_differential_analysis.ipynb THP1_SEC63 CRISPRi perturbation differential expression
bam_index.sh, bam_to_bedgraph.sh,
bedgraph_to_bw.sh THP1_SEC63 BAM -> bedGraph -> bigWig track generation
plotting_marker.ipynb iPSC neuron marker gene plotting
1.0 LPSIFNG_Naive_Kang2019.ipynb Reprocessing of public Kang et al. 2019 microglia RNA-seq
sample_sheet*.csv Sample/condition metadata
microglia_12_states.ipynb,
microglia_12_states_crispri.ipynb Microglia transcriptional-state (12-state) classification,
mapping THP1/HMC3 stimulation states onto human microglial
states from public atlases
See DATA.md for exactly what is and isn't tracked here, and why.
- Recommended/test target: Linux x86_64 (Ubuntu 22.04 LTS or Rocky Linux 8).
- Python: 3.9-3.11; the supplied environment uses Python 3.11.
- R: 4.4 with Bioconductor 3.20-compatible packages.
- Shell: Bash on a POSIX-compatible system.
- Native Windows is not supported. macOS and Windows Subsystem for Linux may work, but have not been validated for the complete ATAC-seq/RNA-seq workflow.
The original notebooks did not record a complete package-version snapshot. The versions in
environment.yml are therefore the maintained reference environment for this
repository, not a claim that every historical notebook was originally run with those exact
versions. A clean-room, end-to-end test is not currently possible because the raw sequencing and
large intermediate data are not included (see DATA.md).
The Conda environment includes the dependencies used across the notebooks and scripts:
- Python 3.11 with JupyterLab 4, pandas 2.x, NumPy 1.x, SciPy 1.x, Matplotlib 3.x, seaborn 0.13, scikit-learn 1.x, and gseapy 1.x.
- R 4.4 with DESeq2, apeglm, IHW, Rsubread, BiocParallel, limma, biomaRt, AnnotationDbi,
GenomicRanges, ChIPseeker, EnhancedVolcano, clusterProfiler, fgsea,
org.Hs.eg.db, tidyverse, cowplot, ggrepel, pheatmap, and msigdbr. - Command-line tools: bedtools 2.31+, samtools 1.18+, Subread/featureCounts 2.0+, GNU parallel,
curl, and UCSC
bedGraphToBigWig.
Some notebooks query Ensembl, Enrichr, MSigDB, or ENCODE and therefore require network access. MSigDB-dependent steps may additionally require acceptance of the relevant data-use terms.
- No GPU is required.
- For a small notebook or test subset: 4 CPU cores and 16 GB RAM are usually sufficient.
- For the full featureCounts/DESeq2 analyses: 8 or more CPU cores and 32-64 GB RAM are recommended.
- Allow at least 100 GB of working disk space in addition to the raw BAM/FASTQ data. Full raw and
intermediate datasets can require several hundred GB;
DATA.mddescribes approximately 700 GB in the original working directories.
The following instructions install the software environment only. The raw and processed datasets
must be obtained separately as described in DATA.md.
-
Install a Conda-compatible distribution such as Miniforge, then clone the repository and enter its directory:
git clone <repository-url> context_ad_multiomics cd context_ad_multiomics
-
Create and activate the reference environment:
conda env create -f environment.yml conda activate context-ad-multiomics
mamba env create -f environment.ymlmay be used instead for faster dependency resolution. -
Register the Python and R kernels so that both are available from Jupyter:
python -m ipykernel install --user \ --name context-ad-multiomics \ --display-name "Python (context-ad-multiomics)" R -e 'IRkernel::installspec(name="context-ad-multiomics-r", displayname="R (context-ad-multiomics)")'
-
Verify the installation:
python -c 'import pandas, numpy, scipy, matplotlib, seaborn, sklearn, gseapy; print("Python dependencies OK")' R -e 'library(DESeq2); library(IHW); library(Rsubread); library(limma); cat("R/Bioconductor dependencies OK\n")' bedtools --version samtools --version featureCounts -v bedGraphToBigWig 2>&1 | head -n 1
-
Start Jupyter from the repository root and select the Python or R kernel required by the notebook:
jupyter lab
Typical installation time: approximately 20-40 minutes on a Linux x86_64 workstation with a broadband connection. Conda dependency solving and Bioconductor package downloads are the main sources of variation; this estimate does not include downloading the study data.
This repository tracks code and sample-sheet metadata only. Raw sequencing reads, peak calls,
differential-analysis result tables, count matrices, and QC reports produced by these pipelines
are not stored here — see DATA.md. Nothing was deleted from the original working
data; it remains on the lab data drive, and will be deposited to SRA/GEO/Zenodo on submission
(accessions TBD).
If you use this code, please cite:
Chen Z, Liu Y, Brown AR, Sestili H, Ramamurthy E, Xiong X, Prokopenko D, Phan BN, Gadey L, Hu P, Tsai LH, Bertram L, Hide W, Tanzi RE, Kellis M, Pfenning AR. Context-dependent regulatory variants in Alzheimer's disease. bioRxiv 2025. https://doi.org/10.1101/2025.07.11.659973
MIT — see LICENSE.