Research registry

Protocols before results

A test gets an ID, controlled files, expected fields, and a named parser environment before it gets a conclusion.

Results pending6

registered test families

External parser runs completed: 0

Research rules

What counts as a result?

CONTROL

One variable changes

Content stays fixed while the file format or layout feature changes.

IDENTITY

The parser is named

A result records the ATS, parser, version when available, date, and configuration limits.

ARTIFACT

Files are retained

Source documents, exports, expected fields, and recovered outputs stay together.

SCOPE

Claims stay narrow

A result for one environment does not become a statement about every ATS.

Protocol registry

Planned test families

Every entry below is a protocol. None is presented as a completed benchmark.

R-01Protocol ready

Multi-column PDF reading order

Changed variable
One-column baseline versus two-column variants
Recorded output
Plain-text order and field adjacency
Open protocol →
R-02Protocol ready

Section-heading recognition

Changed variable
Conventional and custom labels
Recorded output
Section classification and field grouping
Open protocol →
R-03Protocol ready

Icon-only contact information

Changed variable
Visible label, icon plus text, icon only
Recorded output
Email, phone, and URL extraction
Open protocol →
R-05Protocol ready

Unicode punctuation and ligatures

Changed variable
ASCII, Unicode punctuation, accented names, ligatures
Recorded output
Character fidelity and token boundaries
Open protocol →
R-06Protocol ready

Native text versus OCR

Changed variable
Native PDF, scan, and OCR-enhanced scan
Recorded output
Text recovery and field accuracy
Open protocol →

Publication gate

No result without a run record.

  1. 01

    Store the exact source and exported test files.

  2. 02

    Record the parser environment, date, and settings visible to the tester.

  3. 03

    Compare extraction order and fields against the expected-data manifest.

  4. 04

    Repeat failed cases and keep the raw output.

  5. 05

    Publish the narrow result with limitations and no universal pass claim.

References

Sources used on this page

  1. Resume Parser FieldsRChilli Documentation

    Shows the structured fields a commercial parser can return, including experience, education, contact details, and skills.

  2. Unsuccessful resume parseGreenhouse Support

    Documents parse failures, a 2.5 MB parsing limit, and layout-related errors.

  3. Structure of a WordprocessingML documentMicrosoft Learn

    Documents the paragraph, run, and text structure inside DOCX files.

  4. Recognize text in scanned PDFs with AcrobatAdobe Help Center

    Explains that OCR adds a searchable text layer to a scanned PDF.

  5. What you may be missing when you search PDF documentsPDF Association

    Demonstrates how PDF content order can differ from the order visible on the page.

  6. PDF format familyLibrary of Congress

    Notes that scanned-image PDFs do not necessarily support text indexing.