The skills output contains duplicate entries, including repeated Research and multiple Python (Programming Language) items.
What was measured
Noise in output
Does it add incorrect or hallucinated fields?
decisive for this rankingtransformation
Incorrect or hallucinated fields mean the resume parser is not producing reliable structured data, which is central to the job. (3 of 3 judges)
What was given, what came back
Test input: Clean single-column resume — Rugved Nichite · pdf · group: resume-parsing
Input — what we sent

Research media screenshot 202026 05 05 20120443.png
Clean single-column resume — Rugved Nichite
A professionally structured single-column resume for Rugved Nichite, used as the baseline input for parser accuracy across standard resume fields.
Why this input is hard
- · baseline field extraction
- · contact info accuracy
- · work experience parsing
- · education and CGPA extraction
- · skills and certifications extraction
Output — unretouched


Also checked on this input — same tool, 3 other criteria
Accuracy◐ MixedIt identified the degree record but left the CGPA numeric score empty, missing the visible 8.2 value while keeping the unit as CGPA.Accuracy◐ MixedIt split one LinkedIn profile into two website fields and dropped the /in/ path segment, so the URL was not reconstructed as a single complete value.Field coverage✓ WorkedOn the clean resume, the parser covered the standard resume sections expected by the benchmark: identity/contact, work experience, education, skills, and certifications.
Provenance
- Observation
- 33a75e25-fc67-44d4-81d3-e547379126e5
- Evidence run
- cbbef4db-964c-49fa-a57f-a2977822bdfc
- Study
- Parse resumes into structured data using an API
- Research task
- 86b9jm30n
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "affinda",
scenario: "resume-parsing"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on Noise in output
Extracta.ai⚠ StruggledPuts four programming languages — Python, JavaScript, SQL, and Bash — into the Languages field even though the resume does not contain a spoken-languages section.Hireability✗ FailedThe competency extractor introduced noise by labeling all 29 competencies as beginner and including non-skill terms such as Intern, Science, and Framework as competencies.HrFlow✗ FailedIt injects non-skill fragments into the skills list, including 'ml apis', 'lambda', 's3', 'ml', and 'rest apis'.Skima AI✗ FailedThe parser preserved raw bullet/number formatting in responsibilities, leaking duplicated bullet characters into the extracted text.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com