Where your engineering org actually is. What is required to be there. What to measure to prove it. What to stop doing this quarter.
95% of AI pilots never reach production. The reason is almost never the model. It is the skip. Companies jump from L1 to L4 without doing L2 talent or L3 measurement. The dashboard looks good for two quarters, then someone asks what the AI line on the P&L bought us, and no one has an answer.
Below: each level, what is required to be operating at it, the 2 or 3 KPIs that prove you are there, the red flag failure mode, and the next move to climb.
Code touched by AI. Engineers fluent enough to ship with it. AI on the P&L. All three climb. The P&L line lags. That is the L3 trap.
| KPI | Red | Amber | Green |
|---|---|---|---|
| AI-assisted PR rate | <10% | 10 to 30% | >30% |
| Days to first production AI feature | Never set | Slipping | <90 days |
| Paid AI seats deployed | <25% | 25 to 60% | >60% |
"We have 14 pilots running." No, you have 14 ways to stay at L1.
Compliance and security used as the reason for not shipping. Sometimes legitimate. More often a polite stall.
Pick one workflow. Baseline it. Ship in six weeks. Same team owns pilot and production from day one.
| KPI | Red | Amber | Green |
|---|---|---|---|
| % engineers AI-fluent, audited quarterly | <30% | 30 to 70% | >70% |
| Net AI-driven org change, last 12 months | 0% | 5 to 15% | >15% |
| Single named AI-enablement owner | No | In name only | Yes, with QBR |
Layoffs without reskilling. Brand damage now, hiring problem in 18 months. Companies that layoff then rehire spend 1.4 to 2x more over 24 months than companies that reskilled.
Treating AI fluency as a single training course. Real fluency is daily use, not a half-day Zoom.
Stop measuring AI as a cost-saving tool. Start measuring it as a delivery accelerator with a structure that compounds, not one-off layoffs that bleed quality.
| KPI | Red | Amber | Green |
|---|---|---|---|
| AI cost per shipped unit, tracked monthly | Untracked | Tracked, no threshold | Tracked, threshold at QBR |
| AI-attributable margin | Untracked | <2% | >2% |
| AI delivery instability ratio vs non-AI | 2x+ worse | 1.2 to 2x | Within 20% |
Counting Copilot seats as ROI. Seats are an input. Output is what shipped.
50% of leaders want business outcomes. 3% measure them. That gap is where most of the audience sits.
Token costs dropped 98% in 24 months. If your monthly AI bill stayed flat, you are not capturing the deflation. Someone else is.
Build the meta-agent that scores your AI's output. Estimate, per session, how long a human would have taken. That ratio is your L3 proof. Cognition shipped this in 2025. Steal it.
| KPI | Red | Amber | Green |
|---|---|---|---|
| AI-merged PR rate, fully autonomous | <5% | 5 to 30% | >30% |
| Production AI features with CI eval coverage | <40% | 40 to 80% | >80% |
| AI-attributable EBIT, signed by the CFO | 0% | <5% | >5% |
No eval harness. Klarna shipped customer-service AI in Feb 2024 and partially reversed in May 2025. Volume metrics looked great. CSAT and NPS on edge cases collapsed. The team was not measuring quality, only throughput.
Declaring "we are AI-native" in marketing before saying it in the engineering all-hands. If your team does not say it on a Monday, you are not.
Find the next thing AI does that competitors cannot replicate by buying the same models. That is L5. It is a product question, not a delivery question.
| KPI | Red | Amber | Green |
|---|---|---|---|
| % revenue from AI-dependent product capability | <5% | 5 to 20% | >20% |
| AI-native hiring premium vs market | 0% | <10% | >10% |
| Net-new AI-native bets in roadmap, with owners | 0 | 1 to 2 | 3+ |
Calling yourself AI-native because the website says so. Marketing the level you wish you were at is the most common L4 self-deception.
Treating L5 as an end state. Anthropic's internal data shows the complexity and autonomy curves still rising. There is no plateau in 2026.
Three frontiers. Agentic product surface, where your app is the agent. Org as compound interest, where year-two AI-native talent does things year-one talent did not know to ask for. Token deflation as moat-builder, where the workflows that turn cheaper tokens into higher-margin revenue own the next cycle.