Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

context_ad_multiomics

Analysis code for the ATAC-seq and RNA-seq context data characterizing the cell types and stimulation states used to interpret regulatory variant effects in the accompanying Alzheimer's-disease MPRA study (ad_mpra_chen, see below). This covers chromatin accessibility and gene expression profiling of THP1 monocytes/macrophages under cytokine stimulation, primary human macrophages, microglia (HMC3, iPSC-derived, and public single-cell/ex vivo references), iPSC-derived neurons, and public comparison cell lines (K562, HepG2, SK-N-SH) — used to annotate MPRA-tested enhancers with cell-type/state-specific open chromatin and to place THP1/HMC3 stimulation states in the context of human microglial transcriptional states.

This repository holds the analysis notebooks/scripts and sample sheets; the peak calls, differential-analysis result tables, and count matrices they produce are not tracked here — see DATA.md for where to get them.

Relationship to ad_mpra_chen

This repository is a companion to the MPRA analysis repository. Several notebooks there load tables produced by the pipelines in this repository, e.g. notebooks/07_atac_overlap/7.0.atac_annotation_ml_cluster_overlap.ipynb and notebooks/misc/differential_enhancer_ATAC_comparison.ipynb (overlap MPRA allele effects with THP1 differential ATAC peaks). The sequence-based CNN modeling code that consumes this accessibility data lives in ad_mpra_chen/machinelearning_prepost_processing/, alongside the MPRA analysis it directly supports.

Repository layout

atac/
  differential_analysis/
    featurecount_DESEQ2_*.ipynb        featureCounts -> DESeq2/IHW differential peak analysis,
                                        one notebook per comparison/cell-type group
    download_encode_peaks.sh, rename.sh  Fetch/rename public reference peak sets (ENCODE)
    sample_sheet*.csv                  Sample/condition metadata for the notebooks above
  jaccard_index/
    jaccard_index.ipynb                Peak-overlap (Jaccard index) across reference datasets

rna-seq/
  THP1/, THP1_mock_electroporation/, THP1_SEC63/, ipsc_exc_neuron/, Kang2019/
    featurecount_DESEQ2_*.ipynb        featureCounts -> DESeq2/IHW differential expression,
                                        fgsea, per dataset
    CRISPRi_differential_analysis.ipynb  THP1_SEC63 CRISPRi perturbation differential expression
    bam_index.sh, bam_to_bedgraph.sh,
    bedgraph_to_bw.sh                  THP1_SEC63 BAM -> bedGraph -> bigWig track generation
    plotting_marker.ipynb              iPSC neuron marker gene plotting
    1.0 LPSIFNG_Naive_Kang2019.ipynb   Reprocessing of public Kang et al. 2019 microglia RNA-seq
    sample_sheet*.csv                  Sample/condition metadata
  microglia_12_states.ipynb,
  microglia_12_states_crispri.ipynb    Microglia transcriptional-state (12-state) classification,
                                        mapping THP1/HMC3 stimulation states onto human microglial
                                        states from public atlases

See DATA.md for exactly what is and isn't tracked here, and why.

System requirements

Supported platform

  • Recommended/test target: Linux x86_64 (Ubuntu 22.04 LTS or Rocky Linux 8).
  • Python: 3.9-3.11; the supplied environment uses Python 3.11.
  • R: 4.4 with Bioconductor 3.20-compatible packages.
  • Shell: Bash on a POSIX-compatible system.
  • Native Windows is not supported. macOS and Windows Subsystem for Linux may work, but have not been validated for the complete ATAC-seq/RNA-seq workflow.

The original notebooks did not record a complete package-version snapshot. The versions in environment.yml are therefore the maintained reference environment for this repository, not a claim that every historical notebook was originally run with those exact versions. A clean-room, end-to-end test is not currently possible because the raw sequencing and large intermediate data are not included (see DATA.md).

Software dependencies

The Conda environment includes the dependencies used across the notebooks and scripts:

  • Python 3.11 with JupyterLab 4, pandas 2.x, NumPy 1.x, SciPy 1.x, Matplotlib 3.x, seaborn 0.13, scikit-learn 1.x, and gseapy 1.x.
  • R 4.4 with DESeq2, apeglm, IHW, Rsubread, BiocParallel, limma, biomaRt, AnnotationDbi, GenomicRanges, ChIPseeker, EnhancedVolcano, clusterProfiler, fgsea, org.Hs.eg.db, tidyverse, cowplot, ggrepel, pheatmap, and msigdbr.
  • Command-line tools: bedtools 2.31+, samtools 1.18+, Subread/featureCounts 2.0+, GNU parallel, curl, and UCSC bedGraphToBigWig.

Some notebooks query Ensembl, Enrichr, MSigDB, or ENCODE and therefore require network access. MSigDB-dependent steps may additionally require acceptance of the relevant data-use terms.

Hardware

  • No GPU is required.
  • For a small notebook or test subset: 4 CPU cores and 16 GB RAM are usually sufficient.
  • For the full featureCounts/DESeq2 analyses: 8 or more CPU cores and 32-64 GB RAM are recommended.
  • Allow at least 100 GB of working disk space in addition to the raw BAM/FASTQ data. Full raw and intermediate datasets can require several hundred GB; DATA.md describes approximately 700 GB in the original working directories.

Installation

The following instructions install the software environment only. The raw and processed datasets must be obtained separately as described in DATA.md.

  1. Install a Conda-compatible distribution such as Miniforge, then clone the repository and enter its directory:

    git clone <repository-url> context_ad_multiomics
    cd context_ad_multiomics
  2. Create and activate the reference environment:

    conda env create -f environment.yml
    conda activate context-ad-multiomics

    mamba env create -f environment.yml may be used instead for faster dependency resolution.

  3. Register the Python and R kernels so that both are available from Jupyter:

    python -m ipykernel install --user \
      --name context-ad-multiomics \
      --display-name "Python (context-ad-multiomics)"
    R -e 'IRkernel::installspec(name="context-ad-multiomics-r", displayname="R (context-ad-multiomics)")'
  4. Verify the installation:

    python -c 'import pandas, numpy, scipy, matplotlib, seaborn, sklearn, gseapy; print("Python dependencies OK")'
    R -e 'library(DESeq2); library(IHW); library(Rsubread); library(limma); cat("R/Bioconductor dependencies OK\n")'
    bedtools --version
    samtools --version
    featureCounts -v
    bedGraphToBigWig 2>&1 | head -n 1
  5. Start Jupyter from the repository root and select the Python or R kernel required by the notebook:

    jupyter lab

Typical installation time: approximately 20-40 minutes on a Linux x86_64 workstation with a broadband connection. Conda dependency solving and Bioconductor package downloads are the main sources of variation; this estimate does not include downloading the study data.

Data availability

This repository tracks code and sample-sheet metadata only. Raw sequencing reads, peak calls, differential-analysis result tables, count matrices, and QC reports produced by these pipelines are not stored here — see DATA.md. Nothing was deleted from the original working data; it remains on the lab data drive, and will be deposited to SRA/GEO/Zenodo on submission (accessions TBD).

Citation

If you use this code, please cite:

Chen Z, Liu Y, Brown AR, Sestili H, Ramamurthy E, Xiong X, Prokopenko D, Phan BN, Gadey L, Hu P, Tsai LH, Bertram L, Hide W, Tanzi RE, Kellis M, Pfenning AR. Context-dependent regulatory variants in Alzheimer's disease. bioRxiv 2025. https://doi.org/10.1101/2025.07.11.659973

License

MIT — see LICENSE.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages