Benchmarks
Every FitRank model is tested against a simple rules-only ranking on tasks it has never seen. These are the current results, including where it still falls short.
How we measured
40 tasks the model never saw in training, scored across 647 people, October 2026. The data is synthetic, generated with a hidden true fit for every pair so results can be checked exactly.
Synthetic data proves the method works; it doesn't prove accuracy at your company. That's why every pilot measures FitRank against your own managers' decisions before anyone relies on it.
| Measure | FitRank | Rules only |
|---|---|---|
| Best person in the top five | 94.3% | 91.4% |
| Pairs of people ranked in the right order | 88.7% | 86.1% |
| Fit or not-fit decisions correct | 91.3% | 90.4% |
| Calibration error of the fit probabilityLower is better: a 90% score means about 90%. | 2.2% | n/a |
| Correct when the model is at least 90% sureCovers about 70% of people; the rest go to Review. | 98.5% | n/a |
| Scoring time per match run (median)On an ordinary laptop CPU. | 0.94 s | n/a |
| AI cost to score a match runThe decision model runs in-house; no per-call API fees. | $0.00 | n/a |
Companies the model had never seen
Balanced accuracy of fit decisions on day one, before any feedback from that company. Two of these company types were held out of training entirely.
- 85.6%
- IT services company
- 89.5%
- AI lab (never seen in training)
- 91.9%
- QA company (never seen in training)
Where FitRank still falls short
FitRank is better than rules at putting the right person in the top five and at ordering people, but rules-only ranking still places the single best person first more often (74.3% against 68.6%). That's one reason FitRank shortlists rather than assigns, and why managers always see several people with their reasons.
Measure FitRank on your own decisions
Pilot companies get an accuracy review against their managers' real choices before rollout.