本文へ移動
cccskills
無料GitHub で公開

maxquant-proteomics

MaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano plot; GO/pathway enrichment. Use Proteome Discoverer for Thermo-native processing; FragPipe/MSFragger for GPU-accelerated DB search.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md28.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

MaxQuant + Perseus — Proteomics Analysis Pipeline

Overview

MaxQuant is the community-standard software for label-free quantification (LFQ) and SILAC proteomics. It performs database search, protein grouping, and intensity-based quantification from raw LC-MS/MS files, producing proteinGroups.txt as the primary output. Downstream statistical analysis — filtering, normalization, imputation, differential abundance testing, and visualization — is performed in Python using pandas, scipy, and matplotlib/seaborn, mirroring the Perseus workflow in a reproducible scripting environment.

When to Use

  • Performing label-free quantification (LFQ) of proteins across multiple biological conditions — MaxQuant's MaxLFQ algorithm is the community benchmark
  • Running SILAC (stable isotope labeling) experiments with light/heavy or triple-label designs
  • Processing iTRAQ or TMT isobaric labeling experiments via MaxQuant's reporter ion quantification
  • Identifying and quantifying proteins when you need the widely-cited MaxQuant output format (proteinGroups.txt) for comparison with published datasets
  • Performing statistical differential abundance analysis on MaxQuant outputs without installing Perseus (GUI-only, Windows)
  • Generating publication-quality volcano plots and GO enrichment from proteomics data in a reproducible Python workflow
  • Use Proteome Discoverer instead when working with Thermo raw files requiring instrument-native processing or Sequest HT
  • Use FragPipe/MSFragger instead for GPU-accelerated database search (3–10× faster) or when processing DIA (data-independent acquisition) data
  • Use omics-plotting SKILL after differential abundance analysis or GSEA for publication-quality plots

Prerequisites

  • MaxQuant: Windows software; download from https://maxquant.org/ (v2.4+); requires .NET 6 runtime
  • Python packages: pandas, numpy, scipy, matplotlib, seaborn, statsmodels, gseapy
  • Data requirements: Thermo .raw files or mzML-converted files; FASTA protein database (UniProt reviewed + contaminant database)
  • Environment: MaxQuant runs on Windows (GUI or CLI); Python analysis runs cross-platform
pip install pandas numpy scipy matplotlib seaborn statsmodels gseapy
# Install pyMaxQuant for programmatic mqpar.xml configuration
pip install pymaxquant

Pre-flight Interview

Settle these with the user before writing any analysis code.

decisions:
  - id: D1
    param: quantificationStrategy
    kind: required
    source: user
    ask: "How were samples quantified - label-free, SILAC, or isobaric tags?"
    default: null

  - id: D2
    param: proteinSequenceDatabase
    kind: required
    source: user
    ask: "Which organism's protein FASTA, and which release, should spectra be searched against?"
    default: null

  - id: D3
    param: digestionEnzyme
    kind: required
    source: user
    ask: "Which protease was used, and how many missed cleavages should be allowed?"
    default: "trypsin, up to 2 missed cleavages"

  - id: D4
    param: modifications
    kind: required
    source: user
    ask: "Which modifications are fixed, and which variable - and is there an enrichment such as phospho to search for?"
    default: "fixed carbamidomethyl on cysteine; variable oxidation and N-terminal acetylation"

  - id: D5
    param: falseDiscoveryRates
    kind: required
    source: user
    ask: "What false-discovery rate should peptide and protein identifications be held to?"
    default: "1% at both levels"

  - id: D6
    param: matchBetweenRuns
    kind: required
    source: user
    depends_on: [D1]
    ask: "Should identifications be transferred between runs by retention time, recovering more proteins at the cost of some transferred errors?"
    default: "off"

  - id: D7
    param: lfqMinRatioCount
    kind: optional
    source: user
    depends_on: [D1]
    ask: "How many shared peptides must two samples have before their abundances are normalized against each other?"
    default: 2
    skip_if: "labelled quantification"

  - id: D8
    param: missingValueImputation
    kind: required
    source: user
    ask: "How should proteins missing from a sample be treated - left missing, or imputed as below the detection limit?"
    default: "imputed from a downshifted distribution"

  - id: D9
    param: differentialThresholds
    kind: required
    source: user
    ask: "What significance and fold-change cuts define a changed protein?"
    default: "FDR 0.05, log2 fold change 1.0"

D8 is where proteomics differs from RNA-seq and where results most often turn on an unexamined default. A protein absent from every control and present in every treated sample is the strongest possible result or an artefact of detection limits, and the imputation choice decides which one the volcano plot shows. D6 interacts with it: transferred identifications fill in exactly the values that would otherwise be imputed.

Quick Start

import pandas as pd
import numpy as np

# Load MaxQuant output
df = pd.read_csv("combined/txt/proteinGroups.txt", sep="\t", low_memory=False)
print(f"Raw protein groups: {len(df)}")

# Filter contaminants, reverse decoys, only-by-site
mask = (
    (df["Potential contaminant"] != "+") &
    (df["Reverse"] != "+") &
    (df["Only identified by site"] != "+")
)
df = df[mask].copy()
print(f"After filtering: {len(df)} protein groups")

# Extract LFQ intensity columns
lfq_cols = [c for c in df.columns if c.startswith("LFQ intensity ")]
print(f"LFQ columns: {lfq_cols}")

# Log2-transform (0 → NaN)
lfq = df[lfq_cols].replace(0, np.nan)
lfq = np.log2(lfq)
print(f"Valid values per sample:\n{lfq.notna().sum()}")

Workflow

Step 1: Configure MaxQuant Parameters via mqpar.xml

MaxQuant is controlled by an XML parameter file (mqpar.xml). Edit it programmatically to set file paths, enzyme, modifications, and quantification type before running the search.

import xml.etree.ElementTree as ET

def update_mqpar(template_path: str, output_path: str,
                 raw_files: list[str], fasta_path: str,
                 experiment_names: list[str]) -> None:
    """Update mqpar.xml with sample-specific file paths."""
    tree = ET.parse(template_path)
    root = tree.getroot()

    # Set raw file paths
    file_paths_node = root.find(".//filePaths")
    file_paths_node.clear()
    for rf in raw_files:
        elem = ET.SubElement(file_paths_node, "string")
        elem.text = rf

    # Set experiment names (maps files to conditions)
    experiments_node = root.find(".//experiments")
    experiments_node.clear()
    for name in experiment_names:
        elem = ET.SubElement(experiments_node, "string")
        elem.text = name

    # Set FASTA database
    fasta_node = root.find(".//fastaFiles/FastaFileInfo/fastaFilePath")
    fasta_node.text = fasta_path

    tree.write(output_path, xml_declaration=True, encoding="utf-8")
    print(f"Written: {output_path}")

# Example usage
raw_files = [
    r"C:\Data\ctrl_rep1.raw",
    r"C:\Data\ctrl_rep2.raw",
    r"C:\Data\treat_rep1.raw",
    r"C:\Data\treat_rep2.raw",
]
update_mqpar(
    template_path="mqpar_template.xml",
    output_path="mqpar.xml",
    raw_files=raw_files,
    fasta_path=r"C:\Databases\human_uniprot_contaminants.fasta",
    experiment_names=["ctrl", "ctrl", "treat", "treat"],
)

Key mqpar.xml parameters (set in template or edit directly):

<!-- Enzyme and search settings -->
<enzymes>
  <string>Trypsin/P</string>
</enzymes>
<maxMissedCleavages>2</maxMissedCleavages>
<variableModifications>
  <string>Oxidation (M)</string>
  <string>Acetyl (Protein N-term)</string>
</variableModifications>
<fixedModifications>
  <string>Carbamidomethyl (C)</string>
</fixedModifications>

<!-- LFQ settings -->
<lfqMode>1</lfqMode>                     <!-- 1 = LFQ enabled -->
<lfqMinRatioCount>2</lfqMinRatioCount>   <!-- minimum peptides for LFQ -->
<matchBetweenRuns>True</matchBetweenRuns>

<!-- FDR thresholds -->
<peptideFdr>0.01</peptideFdr>
<proteinFdr>0.01</proteinFdr>

Step 2: Run MaxQuant from Command Line (Windows)

MaxQuant can be run headlessly from the Windows command prompt using the bundled MaxQuantCmd.exe.

REM Windows Command Prompt — run MaxQuant with configured mqpar.xml
REM Adjust path to match your MaxQuant installation directory

set MQ_PATH=C:\Program Files\MaxQuant\bin\MaxQuantCmd.exe
set MQPAR=C:\Projects\proteomics\mqpar.xml

"%MQ_PATH%" "%MQPAR%"

REM For specific workflow steps only (useful for reruns):
REM Step IDs: 0=write tables, 1=feature detection, 7=peptide identification
"%MQ_PATH%" "%MQPAR%" --steps 1,7,11
# Cross-platform: run MaxQuant under Wine on Linux/macOS (CI/server use)
wine MaxQuantCmd.exe mqpar.xml

# Monitor progress log
tail -f combined/proc/#runningTimes.txt

Step 3: Load and Filter proteinGroups.txt

Filter out reverse decoys, potential contaminants, and proteins only identified by modification site.

import pandas as pd
import numpy as np

def load_protein_groups(path: str) -> pd.DataFrame:
    """Load MaxQuant proteinGroups.txt with quality filters applied."""
    df = pd.read_csv(path, sep="\t", low_memory=False)
    print(f"Total protein groups: {len(df)}")

    # Remove reverse decoys, contaminants, and only-by-site hits
    n_before = len(df)
    df = df[
        (df.get("Reverse", pd.Series("")) != "+") &
        (df.get("Potential contaminant", pd.Series("")) != "+") &
        (df.get("Only identified by site", pd.Series("")) != "+")
    ].copy()
    print(f"After quality filter: {len(df)} ({n_before - len(df)} removed)")

    # Parse gene names (take first entry for multi-gene groups)
    df["Gene names"] = df["Gene names"].fillna("Unknown").str.split(";").str[0]

    # Set unique index on majority protein ID
    df = df.set_index("Majority protein IDs")
    return df

# Load output
pg = load_protein_groups("combined/txt/proteinGroups.txt")

# Identify LFQ intensity columns
lfq_cols = [c for c in pg.columns if c.startswith("LFQ intensity ")]
print(f"LFQ samples ({len(lfq_cols)}): {lfq_cols}")
# Output: LFQ samples (6): ['LFQ intensity ctrl_1', 'LFQ intensity ctrl_2', ...]

Step 4: Log2 Transform and Median Normalize LFQ Intensities

Replace zero intensities with NaN (missing values in MaxQuant are exported as 0), log2-transform, then apply per-sample median centering.

def prepare_lfq_matrix(df: pd.DataFrame, lfq_cols: list[str]) -> pd.DataFrame:
    """Extract, transform, and normalize LFQ intensity matrix."""
    # Extract and rename columns (strip 'LFQ intensity ' prefix)
    lfq = df[lfq_cols].copy()
    lfq.columns = [c.replace("LFQ intensity ", "") for c in lfq_cols]

    # Replace 0 with NaN (MaxQuant encodes missing as 0)
    lfq = lfq.replace(0, np.nan)

    # Log2 transform
    lfq = np.log2(lfq)

    # Median centering per sample (subtract per-column median of valid values)
    col_medians = lfq.median(axis=0)
    global_median = col_medians.median()
    lfq = lfq.subtract(col_medians, axis=1).add(global_median)

    print(f"Matrix shape: {lfq.shape}")
    print(f"Missing values per sample:\n{lfq.isna().sum()}")
    print(f"Valid values per sample:\n{lfq.notna().sum()}")
    return lfq

lfq_matrix = prepare_lfq_matrix(pg, lfq_cols)
# Matrix shape: (3241, 6)
# Missing values per sample: ctrl_1: 421, ctrl_2: 389, ...

Step 5: Impute Missing Values (MNAR Strategy)

Missing-not-at-random (MNAR) values arise from proteins below the detection limit. Impute from the low end of the observed intensity distribution — the standard Perseus approach.

def impute_mnar(lfq: pd.DataFrame,
                width: float = 0.3,
                downshift: float = 1.8,
                random_state: int = 42) -> pd.DataFrame:
    """
    Impute MNAR missing values from a downshifted Gaussian.

    Parameters
    ----------
    width      : std of imputation distribution (fraction of sample std)
    downshift  : downshift in units of sample std below mean
    random_state : for reproducibility
    """
    rng = np.random.default_rng(random_state)
    lfq_imp = lfq.copy()

    for col in lfq_imp.columns:
        col_data = lfq_imp[col].dropna()
        col_mean = col_data.mean()
        col_std = col_data.std()

        n_missing = lfq_imp[col].isna().sum()
        if n_missing > 0:
            imputed = rng.normal(
                loc=col_mean - downshift * col_std,
                scale=width * col_std,
                size=n_missing,
            )
            lfq_imp.loc[lfq_imp[col].isna(), col] = imputed

    print(f"Imputed {lfq.isna().sum().sum()} missing values")
    return lfq_imp

lfq_imputed = impute_mnar(lfq_matrix)
# Imputed 2847 missing values

Step 6: Statistical Testing — t-test with FDR Correction

Perform two-sample t-tests for each protein between conditions, then apply Benjamini-Hochberg FDR correction.

from scipy import stats
from statsmodels.stats.multitest import multipletests

def differential_abundance(lfq: pd.DataFrame,
                            group_a: list[str],
                            group_b: list[str],
                            alpha: float = 0.05) -> pd.DataFrame:
    """
    Two-sample t-test + BH FDR correction for all proteins.

    Parameters
    ----------
    group_a, group_b : sample name lists for each condition
    alpha : FDR threshold
    """
    results = []
    for protein_id, row in lfq.iterrows():
        a_vals = row[group_a].dropna().values
        b_vals = row[group_b].dropna().values

        if len(a_vals) >= 2 and len(b_vals) >= 2:
            t_stat, p_val = stats.ttest_ind(a_vals, b_vals, equal_var=False)
            log2fc = b_vals.mean() - a_vals.mean()
        else:
            t_stat, p_val, log2fc = np.nan, np.nan, np.nan

        results.append({
            "protein_id": protein_id,
            "log2FC": log2fc,
            "pvalue": p_val,
            "t_stat": t_stat,
        })

    res_df = pd.DataFrame(results).set_index("protein_id")

    # BH FDR correction on valid p-values
    valid = res_df["pvalue"].notna()
    _, padj, _, _ = multipletests(res_df.loc[valid, "pvalue"], method="fdr_bh")
    res_df.loc[valid, "padj"] = padj

    # Add significance flag
    res_df["significant"] = (res_df["padj"] < alpha) & (res_df["pvalue"].notna())
    sig_count = res_df["significant"].sum()
    print(f"Significant proteins (FDR < {alpha}): {sig_count}")

    return res_df.sort_values("padj")

# Define sample groups
group_ctrl  = ["ctrl_1",  "ctrl_2",  "ctrl_3"]
group_treat = ["treat_1", "treat_2", "treat_3"]

results = differential_abundance(lfq_imputed, group_ctrl, group_treat)
print(results[results["significant"]].head(10))
# Significant proteins (FDR < 0.05): 312

Step 7: Volcano Plot Visualization

Plotting: compute the DEG table, then read skills/data-visualization/omics-plotting/SKILL.md and follow its "Volcano" recipe on the exported CSV (→ figures/volcano_plot.png).

# Volcano input: log2FC vs padj + protein labels; render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Volcano" recipe.
volcano_df = results[["log2FC", "padj"]].copy()
volcano_df["gene"] = pg["Gene names"].reindex(results.index)
volcano_df.dropna(subset=["log2FC", "padj"]).to_csv("volcano_input.csv")
print(f"Volcano input: {int(volcano_df['padj'].notna().sum())} proteins -> volcano_input.csv")
# -> figures/volcano_plot.png via omics-plotting (map log2FC -> log2FoldChange)

Step 8: GO/Pathway Enrichment of Significant Proteins

Run over-representation analysis (ORA) on significantly up- and down-regulated proteins using gseapy's Enrichr API.

import gseapy as gp

def run_enrichment(results: pd.DataFrame,
                   gene_names: pd.Series,
                   gene_sets: list[str] | None = None,
                   top_n: int = 20) -> dict[str, pd.DataFrame]:
    """
    ORA enrichment for up- and down-regulated proteins via Enrichr.

    Parameters
    ----------
    gene_sets : Enrichr gene set libraries (default: GO BP + KEGG)
    top_n     : top results to display per direction
    """
    if gene_sets is None:
        gene_sets = ["GO_Biological_Process_2023", "KEGG_2021_Human"]

    results_out = {}
    gene_map = gene_names.reindex(results.index).fillna("Unknown")

    for direction in ("up", "down"):
        sig = results[
            (results["significant"]) &
            (results["log2FC"] > 0 if direction == "up" else results["log2FC"] < 0)
        ]
        gene_list = gene_map.reindex(sig.index).tolist()
        print(f"{direction.capitalize()}-regulated: {len(gene_list)} proteins")

        if len(gene_list) < 5:
            print(f"  Too few proteins for enrichment (n={len(gene_list)}), skipping")
            continue

        enr = gp.enrichr(
            gene_list=gene_list,
            gene_sets=gene_sets,
            organism="Human",
            outdir=f"enrichment_{direction}",
            cutoff=0.05,
        )
        top_results = enr.results.sort_values("Adjusted P-value").head(top_n)
        results_out[direction] = top_results
        print(top_results[["Term", "Adjusted P-value", "Overlap"]].to_string(index=False))

    return results_out

enr_results = run_enrichment(results, pg["Gene names"])

Key Parameters

ParameterDefaultRange / OptionsEffect
matchBetweenRunsFalseTrue / FalseTransfers identifications across runs by retention time matching; increases quantified protein count 10–30%
lfqMinRatioCount21–5Minimum peptide pairs required for LFQ normalization; lower values increase coverage but reduce accuracy
maxMissedCleavages20–4Tryptic missed cleavages allowed; increase for samples with poor digestion
peptideFdr / proteinFdr0.010.001–0.05FDR thresholds for peptide and protein identifications
MNAR downshift1.81.5–2.5Shifts imputation distribution below detection limit in units of column std; larger = more conservative imputation
MNAR width0.30.1–0.5Width of imputed distribution relative to column std
t-test alpha0.050.01–0.1FDR significance threshold for differential abundance
fc_threshold (volcano)1.00.5–2.0log2 fold-change cutoff for "significant" label in volcano plot

Key Concepts

MaxQuant Output Files

FileContent
proteinGroups.txtPrimary output: one row per protein group with LFQ/SILAC intensities, peptide counts, sequence coverage
peptides.txtPeptide-level quantification with charge states and modifications
evidence.txtIndividual MS/MS identifications (one row per peptide-spectrum match)
msms.txtFull MS/MS scan data including fragment ions and scores
summary.txtPer-raw-file statistics: identifications, MS/MS counts, calibration

LFQ Intensity vs. iBAQ

  • LFQ (Label-Free Quantification): MaxLFQ algorithm normalizes intensities across samples based on razor+unique peptide ratios. Use for cross-sample comparisons (fold changes). Stored in LFQ intensity <sample> columns.
  • iBAQ (intensity-Based Absolute Quantification): Divides summed peptide intensities by the number of theoretically observable peptides. Use for estimating copy numbers and comparing absolute abundance between proteins within a sample. Stored in iBAQ column.
  • SILAC ratio: Direct H/L ratio from isotope-labeled pairs. More accurate than LFQ for small fold changes.

Perseus Equivalent Operations in Python

Perseus stepPython equivalent
Filter rows by categorical columndf[df["Reverse"] != "+"]
Replace 0 with NaNdf.replace(0, np.nan)
Log2 transformnp.log2(df)
Median normalizationdf.subtract(df.median()).add(global_median)
MNAR imputation (normal distribution)impute_mnar() function above
Two-sample t-testscipy.stats.ttest_ind() + multipletests()
Volcano plotmatplotlib.pyplot scatter + threshold lines
Hierarchical clusteringseaborn.clustermap()

Common Recipes

Recipe: SILAC Ratio Analysis

When to use: SILAC experiments with H/L or H/M/L labeling instead of LFQ.

import pandas as pd
import numpy as np

# Load proteinGroups.txt for SILAC experiment
df = pd.read_csv("combined/txt/proteinGroups.txt", sep="\t", low_memory=False)

# Filter contaminants and decoys
df = df[(df["Reverse"] != "+") & (df["Potential contaminant"] != "+")].copy()

# Extract H/L ratio columns (log2-transformed)
ratio_cols = [c for c in df.columns if c.startswith("Ratio H/L ") and "normalized" in c.lower()]
if not ratio_cols:
    # Fall back to non-normalized
    ratio_cols = [c for c in df.columns if c.startswith("Ratio H/L")]

print(f"SILAC ratio columns: {ratio_cols}")
ratios = df[ratio_cols].copy().replace(0, np.nan)

# Log2 transform ratios
log2_ratios = np.log2(ratios)
log2_ratios.columns = [c.replace("Ratio H/L normalized ", "") for c in ratio_cols]

# Summary statistics per sample
print(log2_ratios.describe().round(3))

Recipe: Hierarchical Clustering Heatmap

When to use: visualizing patterns across all significant proteins simultaneously. Compute the z-scored matrix below, then read skills/data-visualization/omics-plotting/SKILL.md and follow its "Clustered expression heatmap" recipe (→ figures/heatmap.pdf).

import pandas as pd

def prep_heatmap_matrix(lfq_imputed: pd.DataFrame,
                        results: pd.DataFrame,
                        gene_names: pd.Series,
                        top_n: int = 50,
                        out_path: str = "heatmap_matrix.csv") -> pd.DataFrame:
    """Z-scored expression matrix of the top significant proteins, for the omics-plotting heatmap."""
    sig_proteins = results[results["significant"]].nsmallest(top_n, "padj").index
    heatmap_data = lfq_imputed.loc[sig_proteins].copy()
    heatmap_data.index = gene_names.reindex(sig_proteins).fillna(sig_proteins)
    # Z-score per row
    heatmap_z = heatmap_data.subtract(heatmap_data.mean(axis=1), axis=0).divide(
        heatmap_data.std(axis=1).replace(0, 1), axis=0
    )
    heatmap_z.to_csv(out_path)
    return heatmap_z

mat = prep_heatmap_matrix(lfq_imputed, results, pg["Gene names"])
print(f"Heatmap matrix: {mat.shape} -> heatmap_matrix.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Clustered expression heatmap" recipe -> figures/heatmap.pdf

Recipe: STRING-db Network Enrichment for Significant Proteins

When to use: protein-protein interaction network analysis and enrichment without downloading gene sets locally.

import requests
import pandas as pd

def string_enrichment(gene_list: list[str],
                      species: int = 9606,
                      fdr_threshold: float = 0.05) -> pd.DataFrame:
    """Query STRING /enrichment endpoint for GO/KEGG enrichment."""
    url = "https://string-db.org/api/json/enrichment"
    params = {
        "identifiers": "\r".join(gene_list),
        "species": species,
        "caller_identity": "maxquant_proteomics_skill",
    }
    response = requests.post(url, data=params)
    response.raise_for_status()

    enr_df = pd.DataFrame(response.json())
    if enr_df.empty:
        print("No enrichment results returned")
        return enr_df

    enr_df = enr_df[enr_df["fdr"].astype(float) < fdr_threshold]
    enr_df = enr_df.sort_values("fdr")
    print(f"Enriched terms (FDR < {fdr_threshold}): {len(enr_df)}")
    print(enr_df[["category", "term", "description", "fdr", "number_of_genes"]].head(15).to_string(index=False))
    return enr_df

# Significant up-regulated gene names
up_genes = pg.loc[
    results[(results["significant"]) & (results["log2FC"] > 1)].index, "Gene names"
].tolist()

string_enr = string_enrichment(up_genes)

Recipe: Parse MaxQuant Output with pyMaxQuant

When to use: reading and filtering MaxQuant text files with a higher-level API.

# pyMaxQuant provides typed accessors for MaxQuant output files
# Install: pip install pymaxquant
from maxquant.io import read_protein_groups

# Load with built-in contaminant filtering
pg_clean = read_protein_groups(
    "combined/txt/proteinGroups.txt",
    filter_invalid=True,      # removes reverse, contaminant, only-by-site
)
print(f"Loaded {len(pg_clean)} filtered protein groups")

# Access LFQ columns via helper
lfq_df = pg_clean.filter(like="LFQ intensity")
print(f"LFQ matrix: {lfq_df.shape}")

Expected Outputs

FileDescription
combined/txt/proteinGroups.txtMain MaxQuant output: protein groups with LFQ intensities, peptide counts, unique peptides, iBAQ
combined/txt/peptides.txtPeptide-level quantification with modifications and charge states
combined/txt/summary.txtPer-raw-file QC statistics: identification rates, MS/MS counts
results_differential.csvDifferential abundance table: log2FC, pvalue, padj, significant per protein
volcano_plot.pdfVolcano plot with up/down-regulated proteins colored and top proteins labeled
heatmap.pdfHierarchical clustering heatmap of top significant proteins (Z-score normalized)
enrichment_up/gseapy output directory: GO/KEGG enrichment for up-regulated proteins
enrichment_down/gseapy output directory: GO/KEGG enrichment for down-regulated proteins

Troubleshooting

ProblemCauseSolution
MaxQuant produces 0 protein identificationsWrong FASTA database or enzyme settings; raw file path not foundVerify .raw file paths in mqpar.xml are absolute Windows paths; confirm enzyme matches experiment (Trypsin/P vs Trypsin); check summary.txt for identification rate
All LFQ intensities are 0 after filteringmatchBetweenRuns off + sparse data, or wrong column selectionCheck combined/txt/proteinGroups.txt directly; use pg.filter(like="LFQ intensity") to confirm column names; lower lfqMinRatioCount to 1
Too many missing values after log2 transformInsufficient replicates, inconsistent sample loading, or undetected peptidesEnable matchBetweenRuns; verify equal protein loading (Bradford/BCA); consider stricter valid-value filter (require 3/3 per group) before imputation
Memory error loading proteinGroups.txtFile is large (>500 MB for DDA with many samples)Use pd.read_csv(..., low_memory=False, usecols=[...]) to select only needed columns; or use pd.read_csv(..., chunksize=...)
gseapy Enrichr returns empty resultsGene symbols unrecognized or network timeoutEnsure gene list uses HGNC symbols (not UniProt IDs); check internet connectivity; use gp.enrichr(..., timeout=60)
Volcano plot: all proteins in "ns"FDR threshold too stringent or padj not calculatedVerify multipletests returned valid FDR values; try relaxing alpha to 0.1; check sample group assignments are correct
MaxQuant run hangs at "Feature detection"Low memory (MaxQuant needs 4–8 GB RAM per 3–4 raw files)Process files in smaller batches; increase system RAM; close other applications
Imputation inflates false positivesImputing too aggressively (low downshift)Increase downshift to 2.0–2.5; alternatively, filter to proteins with ≥ 2 valid values per group before testing

References

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

API + Python SDK for ordering cell-free protein expression and binding assays. Submit sequences for expression (10–100 µg), measure binding affinity (KD) against targets, track status, and retrieve results programmatically — no wet-lab setup. Built for ML-guided directed evolution and antibody/nanobody optimization. Requires Adaptyv account and API key.

日本語の概要は準備中です。原文の説明を表示しています。

jaechang-hits/SciAgent-Skills3762026年9月29日 更新

aeon

無料

scikit-learn compatible Python toolkit for time series ML: classify, cluster, regress, segment, transform with 30+ algorithms (ROCKET, InceptionTime, KNN-DTW, HIVE-COTE, WEASEL). Handles panel, multivariate, and unequal-length series. Maintained successor to sktime. Alternatives: sktime (larger ecosystem), tslearn (fewer algorithms), catch22 (features only).

日本語の概要は準備中です。原文の説明を表示しています。

jaechang-hits/SciAgent-Skills3762026年9月29日 更新

AiZynthFinder retrosynthetic route planning (CASP) from AstraZeneca Molecular AI. Monte Carlo tree search guided by a template-based neural expansion policy recursively disconnects a target SMILES until precursors are found in a purchasable stock. Covers config.yml (v4 format), aizynthcli batch screening, the AiZynthFinder/AiZynthExpander Python API, one-step disconnections, custom stocks via smiles2stock, scorers, Retro*/breadth-first/DFPN search alternatives, and reading output.json.gz / trees.json. Use for synthesis route planning, synthesizability screening, and building-block/precursor search. For reaction barriers use neb-irc-activation-energy; for 2D reaction scheme drawing use rdkit-chemdraw-cdxml.

日本語の概要は準備中です。原文の説明を表示しています。

jaechang-hits/SciAgent-Skills3762026年9月29日 更新

Access AlphaFold DB's 200M+ predicted structures by UniProt ID. Download PDB/mmCIF, analyze pLDDT/PAE, bulk-fetch proteomes via Google Cloud. For experimental structures use PDB; for prediction use ColabFold or ESMFold.

日本語の概要は準備中です。原文の説明を表示しています。

jaechang-hits/SciAgent-Skills3762026年9月29日 更新

Annotated matrices for single-cell genomics. Stores X with obs/var metadata, layers, embeddings (obsm/varm), graphs (obsp/varp), uns. Use for .h5ad/.zarr I/O, concatenation, scverse integration. For analysis use scanpy; for probabilistic models use scvi-tools.

日本語の概要は準備中です。原文の説明を表示しています。

jaechang-hits/SciAgent-Skills3762026年9月29日 更新

GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Load matrix, filter by TFs, infer TF-target-importance links, save network. Dask-parallelized to single-cell scale. Core SCENIC component.

日本語の概要は準備中です。原文の説明を表示しています。

jaechang-hits/SciAgent-Skills3762026年9月29日 更新

jaechang-hits のスキルをすべて見る

このスキルの問題を報告する