AI Resume Screener Gave the Same Resume a 66 and a 99
HackerRank open-sourced its large language model-based applicant tracking system last month, and a researcher quickly found a problem: the same resume produced a score as low as 66 and as high as 99 out of 100 when run 100 times through the tool, according to a June 28, 2026 analysis by Dan Kinsky published on his newsletter danunparsed.com.
The tool — HackerRank's hiring-agent repository on GitHub under the interviewstreet account — parses a resume PDF, enriches it with GitHub profile data, and produces a score out of 100 across four categories plus bonus points. The gap in results matters in practice. At an 85-point cutoff, Kinsky found that a qualified candidate submitting the same static resume would fail the automated screen roughly 65 percent of the time purely due to LLM non-determinism, not any change in their qualifications.
Why the scores swing so much
Reviewing Kinsky's breakdown, the rubric allocates points across five buckets: open-source contributions (35 points), personal projects (30 points), work experience (25 points), technical skills (10 points), and bonus points of up to 20. The split is significant — 65 of the first 100 possible points land in categories that require the model to make a subjective judgment call.
When Kinsky isolated results by category, the pattern was clear. Technical skills scoring was nearly deterministic: the resume earned 8 out of 10 in 98 of the 100 runs. Work experience scoring was completely deterministic. Projects and open-source contributions varied wildly run to run, with Kinsky describing the project category as showing "HUGE variation" in qualitative assessments.
He also tested a second model — Gemini — across 50 evaluations. That run produced a tighter distribution (48 to 64) but the non-determinism remained. At a 60-point cutoff on that model, 28 percent of evaluations still ended in a false failure.
Kinsky's conclusion is pointed: the problem is architectural, not a matter of prompt tuning. LLMs are poorly suited to the stable, category-level judgment this rubric demands, and its weighting toward subjective criteria bakes that instability into the score.
The tool being open-source is what made this analysis possible. Most commercial AI-scoring ATS products are black boxes, and our research into the AI-era job search has found candidates have almost no visibility into how automated screens evaluate them.
What this means for job seekers
The practical upshot is uncomfortable: a rejection from an AI-screened pipeline may say less about your resume than it does about which random number the model drew on the day your application was processed.
That does not mean resume quality is irrelevant. What it does mean is that a single rejection from a role you believe you were qualified for should not be read as a verdict. Reapplying to similar postings at the same company, or to near-identical roles elsewhere, is rational — the score may simply land differently the second time.
It also means the most durable resume strategy is to anchor on the categories where AI scoring is stable. Kinsky's data shows that clearly listed technical skills scored nearly identically across all 100 runs. Quantified experience descriptions held up in the deterministic work experience category. The volatile categories were the ones requiring the model to weigh project narratives and infer open-source impact.
Finally, the rubric's 35-point allocation to open-source contributions is worth understanding before you blame a screen-out on your experience. Engineers who built their careers inside company codebases — with little public GitHub history — are systematically disadvantaged by that weighting. It's a rubric design choice, not a measure of competence.
Sources
Dan Kinsky, "HackerRank Open Sourced Its ATS — The Same Resume Scored 66 to 99 Across 100 Runs," danunparsed.com, June 28, 2026 — https://danunparsed.com/p/hackerrank-open-source-ats
interviewstreet (HackerRank), "hiring-agent: AI agent to evaluate and score resumes," GitHub — https://github.com/interviewstreet/hiring-agent
Related Posts

Teen Founders Show Investors Now Chase Skills, Not Degrees

Anduril's Possible $100B Valuation Signals Defense Tech Boom
