The Stress Test: What Clu's Methodology Predicted About IBM's AI Transformation
One of the most important things we can do as a methodology provider is stress-test our own framework publicly. Not to perform confidence, but to demonstrate the kind of epistemic honesty we ask of the organisations we work with.
So we ran a blind retrospective validation of Clu's task-fingerprint methodology against IBM's AI transformation of its HR function. The case is well-documented in the public record: IBM's Chief Human Resources Officer Nickle LaMoreaux disclosed which tasks were automated, which capabilities were protected, and what the organisational outcome was. That gave us a rare opportunity to test our methodology against a known, named, verifiable result.
IBM has not commissioned or endorsed this analysis. We chose the case because it's not client-sensitive, and the documentation is unusually transparent, not because we have a relationship with IBM.
How the validation was structured
We built a corpus of 100 tasks that represent the IBM HR operating model as of May 2023, drawn entirely from public sources prior to the transformation outcomes being announced. Then we fingerprinted all 100 tasks using Clu's methodology, with the fingerprints constructed before reviewing the outcome data. Ten public claims about the transformation were locked as test cases before predictions ran.
The methodology uses a six-part task fingerprint to calculate an augmentability score for each task. The score is a property of the task's cognitive structure, not the job it sits in, which is the core architectural claim we were testing.
What the methodology predicted, and what happened
All ten public claims passed the validation.
The two most significant results: Tasks focused on workforce composition evaluation and productivity analysis, both explicitly named by IBM's CHRO as the preserved, non-automated capabilities, scored 0.14 on augmentability, placing them firmly in the low-automation tier. The methodology identified these as human-critical based solely on cognitive structure, before any outcome data were consulted.
The tier-level gradient was clean: tasks in the highest automation tier (T0) averaged 0.75 augmentability; tasks in the protected human tier (T3) averaged 0.22. The driving-force decomposition correctly distinguished automation-dominant workflow tasks from AI-dominant natural-language tasks, without reference to outcomes.
The methodology didn't predict the future. It correctly characterised the cognitive structure of work in a way that matched what IBM discovered empirically. That's the test that matters.
The honest limitations
We flagged 15 tasks where hindsight contamination was possible, cases where our familiarity with the eventual outcome might have influenced how we characterised the task structure. We kept those fingerprints in the analysis because removing them would have been a form of data cleaning that inflated the result. The validation stands with the contamination flags visible.
We also haven't had the validation independently reviewed yet. The full methodology paper is public precisely so that critics can stress-test our stress test. We'd rather publish with visible limitations than present a curated result that collapses under scrutiny.
A second validation, testing the methodology's ability to predict the failure of an automation programme rather than just its success, is planned.
Why this matters for organisations making AI transformation decisions now
The standard for selecting a workforce AI methodology shouldn't be 'the vendor told us it works.' It should be 'the vendor can show us, in public, what the methodology predicted and what actually happened, with the limitations fully disclosed.'
We've published our work. We've shown the contamination flags. We've described where the methodology is strong and where it needs further testing. That's the epistemic standard we think workforce decision infrastructure should meet, and it's the standard we're holding ourselves to.
Want to see what this looks like for your organisation?
Clu delivers a full structural diagnostic in days, using data you already hold. No integrations. No surveying. No guesswork.
Start making decisions you can stand behind. It's time to get a clu.




