# Digital PDF table extraction benchmark methodology

Fixture ID: `pdf-table-extractor-v1-matrix-20260804`

1. Generate each valid PDF byte stream from the checked-in fixture builder.
2. Run engine `2026.09.19.1` at the fixed clock `2026-09-19T02:46:21.216Z` with its exact source SHA-256.
3. Compare all 16 expected and observed PASS, BLOCK or INDETERMINATE decisions.
4. Require every unsupported parser, geometry, optional-content or revision surface to suppress candidate cell values from JSON, cell-log CSV, per-table CSV ZIP, XLSX, SVG, Markdown and receipt exports.
5. For PASS scenarios, verify table and row counts, formula-safe CSV/XLSX behavior and exact artifact hashes.
6. Recompute the receipt core and every artifact SHA-256 from exact bytes.
7. Separately run the current rectangular spreadsheet readback and independent browser checks below. They must report PASS with no findings; those checks are not included in this 16-scenario benchmark count.
8. Validate representative generated PDFs with Poppler and compare rendered layout separately from parser output.

Reproduce from the source checkout with the configured maintainer Node, spreadsheet and browser dependencies. Serve app/ on http://127.0.0.1:8767 before running the independent browser check. These commands are not commands to run in a browser console:

```text
FASTTOOL_RELEASE_CLOCK=2026-09-19T02:46:21.216Z FASTTOOL_REPORT_PATH=reports/pdf_table_benchmark_build_20260919b.json node scripts/build_pdf_table_benchmark_20260804.mjs
node scripts/pdf_table_rectangular_export_20260919.mjs
FASTTOOL_BASE_URL=http://127.0.0.1:8767 node scripts/pdf_table_independent_redteam_20260919b.mjs
```

This benchmark proves only the checked fixture behavior for the exact source hashes. It does not prove support for every PDF, OCR accuracy, accounting correctness or semantic equivalence with an unseen source.
