Skip to content
Nerdy beta

DX Core 4

Four dimensions that hold each other honest, right up until somebody wants a single number to rank teams with.

The DX Core 4 is Abi Noda, Laura Tacho, Margaret-Anne Storey, Michaela Greiler, and Nicole Forsgren answering the question engineering leaders kept asking them: between DORA, SPACE, and DevEx, which one do we use? Their answer is all three, folded into four dimensions with one key metric each.

The word doing the work is counterbalanced. Speed measured alone produces exactly what you would expect: bigger diff counts, smaller diffs, and engineers who have learned what the dashboard rewards. Pair it with an experience measure and a quality measure and the gaming gets expensive, because moving one number the cheap way visibly damages another.

DimensionKey metricSecondary metrics
SpeedDiffs per engineer, never at the individual levelLead time, deployment frequency, perceived rate of delivery
EffectivenessDeveloper Experience Index (DXI)Time to 10th PR, ease of delivery, regrettable attrition
QualityChange failure rateFailed deployment recovery time, perceived software quality, operational health and security
ImpactPercentage of time spent on new capabilitiesInitiative progress and ROI, revenue per engineer, R&D as a percentage of revenue

Three of these come straight from DORA. Deployment frequency and lead time sit under Speed, change failure rate becomes the Quality key metric, and failed deployment recovery time is mean time to recovery wearing a clearer name. If you already run DORA, you are most of the way to Core 4 and the new work is on the other two dimensions.

The Developer Experience Index

The DXI is the Effectiveness key metric and the piece DX built itself. It is a composite: the average of driver scores across up to 16 drivers of productivity, selected from a larger library that covers code review, build processes, deep work, local development, documentation, requirements quality, technical debt, and on-call experience. DX collects the inputs three ways — system metrics where the data exists, self-reported survey responses where it does not, and experience sampling that catches developers in the flow of work.

Where It Misleads

Counterbalancing is a design property, not a governance guarantee. DX is refreshingly direct that diffs per engineer works only under three preconditions: counterbalance it with an oppositional metric like the DXI, attach no targets or rewards to it, and communicate it carefully as you roll it out. Every one of those is a condition on how your organization behaves, and no platform enforces any of them. The first time a VP puts diffs per engineer next to team names on a quarterly slide, Goodhart’s Law fires and the counterbalance becomes decoration. Core 4 does not solve the measurement problem so much as assume you already solved it.

The index averages away the only signal worth having. Collapsing 16 drivers into one number reproduces the story-point problem at organizational scale. A three-point DXI gain tells you nothing you can act on. “Deployment confidence dropped and onboarding time doubled” tells you where to go. The drivers are where the insight lives; the index is a trend line and a conversation opener. Report the drivers, and never put the composite in a cross-team comparison.

Three and a half of the four dimensions never leave the engine room. In Impact-Outcome terms, Speed and Quality are output and flow measures, and Effectiveness measures the conditions activity happens under. Only Impact reaches toward the business, and its key metric is an allocation: percentage of time spent on new capabilities tells you where the hours went, not whether the capabilities changed anyone’s behavior. You can move all four dimensions and ship a quarter of features nobody wanted. That is the Output-Activity Trap with better instrumentation.

The Impact dimension argues against work you need. Optimizing the share of time spent on new capabilities pushes directly against KTLO/BAU that is load-bearing, and our position there is that healthy maintenance runs 15-25% and the danger is invisibility, not the percentage. Revenue per engineer and R&D as a percentage of revenue are ratios a CFO can improve by cutting the denominator, which is a fact worth knowing before you put them on a slide.

A perception survey measures what people are willing to say. Response rate and psychological safety bound the validity of every self-reported driver. In a low-trust organization the DXI is measuring trust, and a pathological Westrum culture will return a flattering score right up until the attrition numbers arrive.

How We Use It

We run a measurement engagement on DX’s platform, so treat everything above as an assessment written by someone with a stake in the answer. Here is the falsifier, stated plainly: if the readout says your constraint is decision latency, unclear ownership, or a roadmap nobody believes, then the honest recommendation is a Gemba walk and an org-design conversation, not a platform renewal.

Used well, Core 4 is a discovery instrument rather than a scorecard. It is a fast, rigorous way to convert “everything feels slow” into a ranked list of validated hunches, which is precisely the input the Discover-Option-Action Cycle needs. That is a defensible claim. “Measuring developer productivity” is not, and the framework’s own authors spent two papers explaining why.

Resources