The capability survey
Written by a person. Last read by a person on 2026-09-07, 21 days ago. Its facts were checked by the eval suite on 2026-09-28.
You want the evidence behind the diagnosis, or you are judging whether the survey was run honestly.
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. 1,237 Halden Systems staff answered, a 41% response rate. The 78 and the 31 that every document quotes are not survey items. The 78 is a completion count from the learning platform. The 31 is the pass rate on a scored check of the 1,150 who work in engineering systems daily, described in the capability check below.
What this survey measured is what sits underneath that gap. The items least convenient for this program are in the results, because a survey where every item supports the proposal is a survey nobody ran.
The instrument
| Fielded | 4 to 22 May 2026, 3 weeks |
| Population | 3,000, all non-technical staff |
| Responses | 1,237 |
| Response rate | 41% |
| Languages | 5, the working languages of the 9 sites |
| Design | 14 Likert items, 4 behavioral frequency items, 3 free-text |
Two design decisions are worth stating, because both change what the numbers mean.
We surveyed all 3,000, not the 1,150 who work in engineering systems daily. Sampling only the inner population would have measured the program's convenience rather than the company's problem, and the assistant went to everybody regardless of whether their work touches a repository.
Free text came last. A respondent who has already answered 18 structured items has been reminded what the subject is, and writes something specific instead of something general.
The headline
| Item | Result |
|---|---|
| Worry about breaking something that cannot be undone | 64% |
| Times a week they ask a technical colleague for help, median | 2.3 |
The first row is not timidity. Somebody who cannot predict what a step will do is correct to fear an irreversible one, and the second row is the cost of that fear landing on an engineer. The 47 points between the 78 and the 31 is the finding the program exists to answer; the instrument that produced the 31 is next.
The capability check
The 31 is not an agreement percentage. It is the share of the inner population who could predict what one of our systems would do, on a scored check. It is the measure the program will be held to. Everything here is repeatable. The condition under which we would stop trusting it is stated in advance.
The six items
Each item shows a real state from one of our 4 engineering systems and asks what a named next action will do, and why:
- a branch that has diverged
- a preview that renders correctly against a build that fails
- an automation that read the wrong field
- a permissions error on an action that cannot be undone
- an assistant answer that is wrong and sounds certain
- a pipeline that failed on something other than the code
Scored against one correct outcome plus the reason, with partial credit for a right answer reached by wrong reasoning. Pass is 5 of 6. The reason is scored because a person who predicts correctly without being able to say why cannot do it on the seventh system, and the seventh system is already coming.
The distribution
Fielded over 2 weeks to the 1,150; 712 answered, a 62 percent response rate. Percentages are of the 712 respondents, not of the 1,150.
| Items predicted correctly | Share of respondents |
|---|---|
| 5 or 6 | 31 percent |
| 3 or 4 | 29 percent |
| 2 or fewer | 40 percent |
The middle row is the one to watch over time. It is the population that moves first, and it is where the first 2 quarters of the program should show up if the design is right.
Why the comparison is imperfect, stated in full
Different instruments measuring different things. Completion is an event. Prediction is a performance. Nothing makes the two arithmetically comparable, and we are not claiming a difference of means. We are claiming a direction.
Different denominators. The 78 is a census of 1,150. The 31 is 712 volunteers. Re-weighting the result by site and job family to the known shape of the population moved it by under 2 points. We report the unweighted figure, because the adjustment is not what carries the argument and a reader should not have to take a correction on trust.
Non-response runs in the flattering direction. A person who suspects they would do badly on a capability check is the least likely to take one. Completion is inflated for the mirror-image reason. Both errors push the gap down, which is why we describe 47 as a floor and claim nothing stronger for it.
What the instrument becomes
- An item bank of 30, fielded as parallel 6-item forms, so the check cannot be passed by having taken it before.
- Items decay with the systems. About 12 percent of the content library is wrong at any moment; items about those systems are wrong at the same rate, and the teams that own each system retire them on the same cadence they carry their release notes.
- The bank is published internally. Nobody is being trick-tested. A person who studies the bank and can then predict what the systems do has done the thing we wanted.
- Quarterly, census-attempted, with manager time protected, reporting the response rate alongside the result every time.
- Reported by site first. Proximity to help predicts capability better than job family does, and a single company number hides the 5 sites the program exists for.
- The half-time analyst owns it, including the judgment about when the number has stopped meaning anything.
What this number will never be asked to do
- Carry attribution. A rising pass rate is evidence the program is working. It is not proof that nothing else was happening.
- Benchmark against other companies. No comparable instrument exists outside Halden, and any figure offered for that purpose was invented by somebody.
- Appear in the revenue case. Revenue attribution is the weakest evidence this function offers, and mixing a capability measure into it is how a good number gets spent on a bad argument.
The assistant
| Item | Result |
|---|---|
| Have used it at least once | 71% |
| Use it weekly or more | 44% |
| Would put their name on its output | 38% |
| Have had an answer that was wrong and sounded confident | 57% |
| Know which of its answers need verifying | 22% |
The drop from 71% to 44% is where the rollout actually stands, and neither figure is the one that matters. Judging when the output is wrong is the skill, and 22% report having it.
What people will not say to a manager
| Item | Result |
|---|---|
| Believe the assistant will reduce headcount in their area | 46% |
| Would say that to their manager | 12% |
| Have avoided asking a question because of how it would look | 51% |
| Believe becoming skilled with it would prove they are replaceable | 39% |
| Believe leadership has been straight about it | 23% |
The 34-point gap between the first 2 rows is why leadership's read of this is wrong. Every conversation on the subject has been with the 12%.
The last row sits against a reassurance we already gave. We told people nobody is losing their job, and under a quarter believe we have been straight with them. Those are the same fact.
Time, and what gets in the way
| Item | Result |
|---|---|
| Report more simultaneous tool or process changes than they can absorb | 68% |
| Have no learning time that is not taken from delivery | 74% |
| Do this learning outside working hours | 41% |
The third row is an equity problem before it is a capability problem. It selects for people without caring responsibilities, and any plan that assumes it will continue is choosing that.
Incentives, and what people would actually do
| Item | Result |
|---|---|
| Believe effort to build these skills is noticed | 19% |
| Would answer colleagues' questions if that counted for something | 61% |
| Would be motivated by a reward for completing training | 16% |
| Would accept being publicly identified as learning this | 27% |
Rows 2 and 3 point in opposite directions and settle the incentive design. Paying for completion motivates 16%, and completion is already at 78%. Recognizing contribution reaches 61%, and contribution is the thing we have none of.
Row 4 constrains how. Any scheme that requires somebody to be visibly a learner reaches about a quarter of the population, and loses the reset expert entirely.
By site and by language
| Item | Result |
|---|---|
| Confidence gap, sites with a local expert against sites without | 22 points |
| Response rate at the smallest South American site | 19% |
| Confidence gap for respondents working in English as a second language | 14 points |
The 22-point gap is wider than the gap between job families, which is the finding that makes proximity to help a better predictor of capability than role.
The 19% is reported rather than hidden. It is the site we understand least, and it is one of the 5 with nobody to ask, so the true picture there is probably worse than these results suggest.
The results that are inconvenient for this program
Kept because removing them would make the rest untrustworthy.
| Item | Result |
|---|---|
| Rated previous training as useful | 29% |
| Believe more training would help | 34% |
| Would attend an optional live session | 18% |
29% found our previous training useful. That is this function's own report card, and it is the strongest argument against giving this function more money.
Only a third think more training is the answer. They are largely right, which is why the program is built around local help, contribution and time rather than courses.
18% would attend an optional session. Any plan resting on voluntary synchronous attendance is refuted by its own survey before anybody costs it.
Free text, coded
3 free-text items, 1,237 responses, 7 themes. Counts are responses coded to each theme, and a response can carry more than one.
| Theme | Coded | What it sounds like |
|---|---|---|
| I was good at my job and now I am not | 218 | competence reset |
| Elaborate workarounds nobody can see | 174 | coping machinery |
| I might break something and not be able to undo it | 161 | fear of permanence |
| It was confidently wrong and I stopped | 143 | burned once |
| If I get good at it, I have proved they do not need me | 189 | the replacement bind |
| This is the fourth thing this year | 152 | no time to absorb |
| The person who could tell me is asleep | 128 | no local help |
"I have been a writer for fifteen years. I was the person people asked. Now I copy commands from a document and hope."
The largest theme is not about tools. It is about having been competent and no longer being so, and a program that treats it as a skills gap will be answering a different question than the one people are asking.