An employee messages HR to say their leave balance is wrong. The screen says eleven days and they are certain it should be fourteen. It is a small, ordinary Tuesday-afternoon exchange, and it contains the entire argument about machine learning in HR software.
If the balance came from a rule, there is an answer. Your policy accrues at one and a quarter days a month, you joined in March, three days were taken in June and approved by your manager on the ninth, here is the ledger, and if any line of that is wrong we will fix the line. If the balance came from a model, there is no answer. There is a number, and the system’s confidence in it.
Nobody is currently proposing to predict leave balances. That is the point worth holding onto: the line does not get crossed in one obvious step. It erodes from the other end, where a vendor adds a helpful summary, then a suggested rating, then a flag on a name, and at no single release does anyone stand up and say we have started guessing about people.
The useful half is genuinely useful
It is worth being precise about what works, because an argument that treats every model as a threat gets ignored by the people who most need to hear it.
A model reading an uploaded contract and filling in the start date, the notice period and the job title is doing real work. So is finding the record when somebody types half a name and the wrong spelling of a city. So is compressing a forty-message approval thread into five lines for the manager who has to decide something today. These are extraction, retrieval and summarisation, and they share one structural property: the output sits next to its source, a human can check it in seconds, and correcting it costs nothing but the correction.
Notice what none of them do. They do not decide. They put a draft in front of a person who decides, and the record afterwards shows what the person chose, not what the machine suggested.
Two questions that draw the line
There are two tests, and a system that fails either one should not be producing an outcome about a person.
The first is determinism. Does the same input always produce the same output? A rule does. Change nothing about a person and their entitlement does not move. A model’s output can move because the model was retrained, because a threshold was tuned, or because the population around the person changed. When an outcome moves while the person stands still, that outcome was never a fact about them.
The second is contestability. Can the person affected be told why, in terms they can check and challenge? This is the harder test and the more important one. A rule gives a reason with parts, and each part can be disputed separately: I did not take that day, my joining date is wrong, that policy version was not in force yet. A score gives a number. The honest explanation of a number is a description of how the model works, which is not a reason and cannot be argued with. The best a vendor can offer is a list of the features that pushed the number around, which tells the employee what correlated, never what was true about them.
Where it must not go
Four things fail both tests, and the reasons are worth stating rather than assumed.
Scoring a person. Attrition risk, engagement, potential, whatever it is called this year. A probability attached to a name changes how that name is read in every meeting afterwards, and it is unfalsifiable: if the person stays, the score was a useful warning, and if they leave, the model was right. Nothing about the person was ever established.
Ranking employees against each other. A rank is not a measurement, it is a forced ordering, and it produces a loser at every level no matter how good the level is. When the ordering is generated rather than argued, no manager owns it and no employee can appeal it.
Screening candidates out. A model trained on past hiring reproduces past hiring. The candidates removed before a human looked never learn why, and neither does the employer, which means an unlawful pattern can run for a year without anyone being in a position to notice it.
Any number that lands on a payslip or an entitlement. A leave balance, a notice period, a settlement figure, a gratuity, a pro-rated deduction. These are legal quantities with a right answer, and their right answer comes from a policy document and a calendar. A model that approximates them is wrong in a way nobody can see until it matters.
A rule can be audited, a guess cannot be explained to a tribunal
There is a practical version of this argument that carries more weight than the ethical one, and it is worth leading with in front of a board.
When an employment decision is challenged, at a tribunal, in a regulator’s questionnaire, or in the discovery attached to a dispute, the employer has to produce the reason. Not the intention, the reason: the rule that was applied, the version of the policy in force on the day, the inputs, and the record of who decided. A rule survives this. A rule has an effective date, a version, and a stored trail of what it produced, so a decision from three years ago can be reproduced exactly.
A model does not survive it. The version that produced the score has probably been retrained since. The training data has moved. The threshold was adjusted in a release nobody documented as a policy change, because it was not filed as one. Even where the score itself was logged, the reasoning behind it was never a thing that existed in a form anybody can read back.
And the burden lands on the employer rather than the vendor. Regulators are converging on treating employment software as high-risk precisely because the consequences are borne by the person with the least ability to see the mechanism, and the duties attach to whoever made the decision. The guide on HR software without AI surveillance sets out how that regulatory picture is developing and what to ask a vendor before you sign.
Where we stand
Everything Capstan determines comes from a rule you can read: an entitlement, an access decision, an approval route, a line in a generated letter. Ask why a leave request routed to a second approver and the answer is the workflow definition and the version of it that was live that day, not a probability.
Payroll numbers are not guessed here either, for a different reason. Capstan computes no statutory pay at all and holds no rate, slab or formula anywhere. It compiles the inputs, hands a documented export to your payroll partner, and files their computed results back onto the record, which means the numbers on a payslip are the partner’s arithmetic rather than anybody’s approximation.
There is no AI capability built in the product, no AI credential is required or read by any production code path, and workspace data is never sent to an outside model or used to train one. The security page states that with the mechanisms underneath it rather than as a slogan, and the one AI vendor on the register is published as admitted but not engaged, because a permission nobody has used is a different thing from a vendor that holds your data.
Here is the honest limit. Résumé parsing, one of the uses this post calls legitimate, is not built in the recruitment module. It is a recorded deferral rather than a refusal, and if it ships, it ships as a draft a human corrects, with the extracted fields visible beside the document they came from. The refusal that does not move is the other one: nothing here will score, rank or screen a person. If that ever changes, the manifesto changes first, in public, before the feature exists.