Reports
speccheck reports are designed to be both human-reviewable and easy to archive.
The important rule is that collected per-sample CSVs are small and reproducible;
large upstream files stay in the workflow output area.
Files written by collect
For each sample, collect writes a concise CSV containing:
- parsed metrics from recognised tools;
PASS/WARN/FAILstatus columns generated from the active criteria;- provenance columns such as
speccheck_version,speccheck_criteria_sha256, andspeccheck_input_file_count; - metadata columns when
--metadatais supplied.
For some wide upstream outputs, a detailed.*.csv companion may also be written.
summary ignores these detailed companions by default so the merged cohort
report remains stable.
Files written by summary
summary writes:
report.csv: merged concise cohort table;report.full.csv: human-ordered parser, compatibility, metadata, and provenance table;report.html: self-contained HTML review report when--plotis enabled;report.xlsx: optional workbook when--xlsx-outputis supplied.
When merging inputs, summary rejects duplicate or missing sample IDs instead
of silently overwriting samples. Sample IDs are read as strings, so identifiers
such as 00123 retain their leading zeroes. Blank IDs, surrounding whitespace,
embedded-ID mismatches, and conflicting canonical/alias metric values stop the
merge with an actionable error.
Key report columns
Start review with these columns:
| Column | Meaning |
|---|---|
overall_qc |
Worst current evaluated status across Speccheck and optional QualiBact QC. |
speccheck_qc |
Native Speccheck status, including criteria not derived from QualiBact. |
qualibact_qc |
Current QualiBact evaluation when --qualibact-compat is enabled. |
historical_qualibact_qc |
Imported comparison label; never used as the current verdict. |
reason_summary |
Compact explanation of failures, warnings, and missing checks. |
speccheck_warning_count |
Number of warning-level criteria triggered. |
speccheck_failure_count |
Number of failure-level criteria triggered. |
speccheck_not_evaluated_count |
Number of expected metrics missing from detected parser outputs. |
species / species_confidence |
Resolved species information where parser outputs provide it. |
top_abundance_percent |
Highest Sylph taxonomic abundance, always on a 0–100 scale. |
report_schema_version |
Version of the stable cohort-report column contract (1.0). |
Tool-specific *.status columns use the same vocabulary:
PASS: evaluated and passed;WARN: evaluated and triggered at least one warning criterion;FAIL: evaluated and triggered at least one failure criterion;NOT_EVALUATED: expected evidence was missing.
CSV QC fields never use booleans or PASSED/FAILED. Legacy input *.check
columns are normalised to *.status in report.full.csv. Missing evidence can
remain NOT_EVALUATED at metric level, while non-strict missing evidence makes
the sample-level Speccheck result WARN.
The concise report is ordered as identifiers, current statuses and reasons,
species assignment, core assembly metrics, taxonomic abundance, then threshold
source. Its 19 version-1.0 columns are always present; optional unavailable
values are blank rather than removing columns. report.full.csv begins with
exactly the same columns, followed by
tool-level and metric statuses, tool metrics grouped by parser, compatibility
details, provenance, and metadata. Proven-equivalent aliases such as
Checkm.GC/Checkm.GC_Content are collapsed to the parser-native canonical
column in the cohort report, but only after verifying that populated aliases
agree. Raw per-sample detailed.*.csv files retain the original fields.
The exact concise order is:
sample_id, overall_qc, speccheck_qc, qualibact_qc,
historical_qualibact_qc, reason_summary, species, species_confidence,
n50, contigs, genome_size, gc_percent, completeness, contamination,
depth, top_species, top_abundance_percent, threshold_source,
report_schema_version
QualiBact compatibility columns
When --qualibact-compat is used, summary adds a species-aware evaluation
from the packaged, release-pinned QualiBact snapshot, including:
qualibact_qc;qualibact_compat_reasons;qualibact_compat_source.
The packaged snapshot covers 317 QualiBact-indexed species and records the preferred scheme used for each species. This is threshold-behaviour alignment, not a claim that Speccheck reproduces the entire QualiBact web application.
top_abundance_percent is always expressed on a 0–100 percentage scale.
HTML report
The HTML report includes:
- cohort-level PASS/WARN/FAIL counts;
- clickable KPI filters and a review queue that shows exceptions by default;
- a focused sample-detail dialog containing key metrics, tool statuses, and failed checks;
- top warning and failure reasons;
- collapsed cohort-level metric summaries with all-species and per-species views;
- collapsed per-tool diagnostics with exception-only result tables;
- lightweight inline SVG charts with hover values and sample-detail links;
- a collapsible data and provenance area for citations and the full-width table;
- persistent desktop navigation and a compact mobile jump menu.
Generate an HTML report with:
speccheck summary qc_collect \
--output qc_report \
--plot \
--qualifyr-style \
--xlsx-output qc_report/report.xlsx
Select FAIL, WARN, NOT_EVALUATED, PASS, or all samples using the review tabs. Interactive tables can be sorted by clicking headers, filtered with the search box, and paged in groups of at most 25 rows. Opening one tool diagnostic closes the others, so charts and secondary evidence do not dominate the page. Chart points use the same PASS/WARN/FAIL/NOT_EVALUATED colours as the rest of the report; selecting a point opens that sample's detail view. The HTML contains its CSS, JavaScript and SVG charts, so it can be stored with a release artifact or shared for review without a server.
For large runs, start with:
- select the FAIL or WARN KPI, or use the default needs-review queue;
- open an individual sample row to inspect its key evidence;
- open a tool diagnostic only when the reason needs investigation;
- use cohort metrics for distribution-level checks;
- open data and provenance only when you need exports, citations, or raw parser columns.
Example report generation
Minimal fixture-based examples:
pixi run python scripts/generate_qualibact_example_reports.py
Real GHRU-derived panel:
pixi run python scripts/build_ghru_ecoli_panel_report.py \
.demo_work/ghru_ecoli_panel/triplet/output \
--metadata .demo_work/ghru_ecoli_panel/triplet/metadata.csv \
--work-dir .demo_work/ghru_ecoli_panel/triplet/work
100-sample case-study assets:
pixi run python scripts/create_real_run_100_assets.py
Docker usage
Typical Docker summary run:
docker run --rm \
-v $(pwd):/data \
-v $(pwd)/output:/output \
happykhan/speccheck \
summary /data --output /output --plot