Skip to content
FitRank

Benchmarks

Every FitRank model is tested against a simple rules-only ranking on tasks it has never seen. These are the current results, including where it still falls short.

How we measured

40 tasks the model never saw in training, scored across 647 people, October 2026. The data is synthetic, generated with a hidden true fit for every pair so results can be checked exactly.

Synthetic data proves the method works; it doesn't prove accuracy at your company. That's why every pilot measures FitRank against your own managers' decisions before anyone relies on it.

FitRank compared with a rules-only ranking
MeasureFitRankRules only
Best person in the top five94.3%91.4%
Pairs of people ranked in the right order88.7%86.1%
Fit or not-fit decisions correct91.3%90.4%
Calibration error of the fit probabilityLower is better: a 90% score means about 90%.2.2%n/a
Correct when the model is at least 90% sureCovers about 70% of people; the rest go to Review.98.5%n/a
Scoring time per match run (median)On an ordinary laptop CPU.0.94 sn/a
AI cost to score a match runThe decision model runs in-house; no per-call API fees.$0.00n/a

Companies the model had never seen

Balanced accuracy of fit decisions on day one, before any feedback from that company. Two of these company types were held out of training entirely.

85.6%
IT services company
89.5%
AI lab (never seen in training)
91.9%
QA company (never seen in training)

Where FitRank still falls short

FitRank is better than rules at putting the right person in the top five and at ordering people, but rules-only ranking still places the single best person first more often (74.3% against 68.6%). That's one reason FitRank shortlists rather than assigns, and why managers always see several people with their reasons.

Measure FitRank on your own decisions

Pilot companies get an accuracy review against their managers' real choices before rollout.