Agentic Engineering

Which systems should agents rebuild, and is adoption real?

The two questions every engineering leader is being asked about AI. Qualimetry answers both from your own estate: a disposition for every system based on pressure and feasibility, and an adoption measure built from signals your organisation controls.

Why this is answerable at all

Handing a system to agents is a governance decision before it is a technical one

Readiness scoring is not an abstract exercise. A system is only safe to hand over once the rules that govern it are somewhere an agent can read, and once what comes back can be measured against those same rules. Every input to the score below depends on the rest of the platform.

Without those four, an agentic readiness score is a guess about a system nobody can govern. With them, it is a decision you can defend.

Agentic Candidates

Every system gets a disposition, not a hunch

Two axes decide it. Replacement pressure is what the business is paying to keep the system as it is. Replacement feasibility is how safely agents could take it on. Where a system lands gives you the verdict, and the evidence behind it.

Agentic disposition on two axes

Every system is placed on two axes. Replacement pressure is how much the business is paying to keep it as it is. Replacement feasibility is how safely agents could take it on, given test coverage, interface clarity, specification stability and the risk of cutting over. The four resulting verdicts are Rebuild, Contain, AI-maintain and Maintain.

  • Rebuild Pressure: High Feasibility: High It costs you to keep, and agents can safely take it on. This is where an agentic programme pays back fastest.
  • Contain Pressure: High Feasibility: Low It costs you to keep, but it is not safe to hand over yet. Limit the blast radius and fix the things blocking feasibility.
  • AI-maintain Pressure: Low Feasibility: High It is not hurting, and agents could work on it comfortably. Let agents carry the routine maintenance.
  • Maintain Pressure: Low Feasibility: Low It works, it is cheap to keep, and there is no case for disturbing it. Leave it alone.
The verdict is the argument, not the answer. Every cell links to the evidence that put a system there.
Readiness

Banded, so the conversation is short

Readiness is a weighted average of the criteria that actually determine whether agents can work safely on a system. A criterion with nothing to measure scores neutral, never zero, so an unmeasured system is not falsely condemned.

Strong

Agents can start here now.

After prep

Fix the blockers first, usually coverage or interface clarity.

Not ready

Too much would have to change before this is safe.

Not a fit

Agentic work is not the right answer for this system.

Your priorities, your weights

The ranking should reflect what you are optimising for

Start from one of four profiles, then adjust the weights directly. Nothing is hidden in a black box, and the ranking recalculates against the criteria below.

Balanced Effort reduction Modernisation Risk reduction

What gets scored

Codebase size, test coverage, active churn, contributor concentration and bus risk, tech-stack obsolescence, stack standardness, maintenance burden, interface clarity, specification stability, verification readiness, business risk during cutover, dependency exposure, compliance, lifecycle, and end-of-life or known-exploited exposure.

Programmes

A programme opens itself, and closes itself

A project starts an agentic programme on its first agent-attributed pull request, and the programme completes after thirty days without another. You do not have to declare a programme, and nobody has to remember to close one.

Agentic Adoption

Signals, not guesswork

Qualimetry does not infer authorship from how code is written. Stylometry is a deliberate non-goal, because a number produced that way cannot be defended when someone senior challenges it. Adoption is measured from two independent signals your organisation controls.

How agentic work is attributed

Two independent signal sources are used and are never blended: work observed to come from a recognised agent account, and work that follows a convention your organisation agreed, such as a branch prefix, a title prefix or a pull request label. Dependency and CI bots are removed from both the numerator and the denominator. What remains is reported as a floor, because work that follows no convention and comes from no recognised account cannot be attributed even if an agent produced it. Writing style is never used to guess.

  • Observed agent The pull request was authored by an account you recognise as an agent ClaudeCursorCodexGitHub CopilotBlitzyYour own accounts
  • Agreed convention The work follows a marker your teams agreed to use Branch prefixTitle prefixPull request label

The two lanes are counted separately and never blended.

Dependency and CI bots are removed Excluded from the numerator and the denominator, so they cannot inflate or dilute the rate
Attributed agentic work Reported as a floor, never as the true rate Attributed Not attributable, may still be agentic
There is no stylometry here. Nothing infers authorship from how the code is written, which is why the number can be trusted in the direction it claims.

Two sources, never blended

An observed agent is a pull request authored by an account you recognise as an agent. Claude, Cursor, Codex, GitHub Copilot and Blitzy are recognised out of the box, and you can add your own accounts and markers.

A convention is a marker your teams agreed to use: a branch prefix, a title prefix, or a pull request label. It catches agentic work that runs under a human account.

  • The two are reported separately, because mixing them would hide which one you can actually rely on.
  • Dependency and CI bots are excluded from the numerator and the denominator.
  • Work that follows no convention appears as untracked rather than being quietly dropped.
  • The estate-wide split stays hidden until convention coverage is high enough to mean anything.
What you see

Trend, coverage, tool mix, governance and outcomes

Adoption trend by source

How attributed agentic work has moved over time, split by the signal that identified it.

Attribution coverage

How much of your estate is even capable of being attributed, which is the honest health check on the number itself.

Adoption by category

Where agentic work concentrates across your own business structure, whether that is domain, journey or owning team.

Tool mix by source

Which agents are actually being used, with convention-only work shown as untracked rather than misattributed.

Change governance

Volume is the easy question. The one that matters to a risk committee is whether agentic changes are being reviewed at all.

  • Approval coverage across the estate.
  • Unreviewed merge rate, with an optional agent overlay.
  • AI review coverage, so you can see where the reviewer is not running.
Honest by design

Every adoption figure is a floor, not the true rate

Work that comes from no recognised account and follows no agreed convention cannot be attributed, even when an agent produced it. Qualimetry says so, in the product, next to the number. A measure you can defend when challenged is worth more than a measure that flatters the programme.

Questions

Agentic engineering, answered

How do you detect AI-written code?
We do not detect it, we attribute it. Two independent signals are used: pull requests authored by an account you recognise as an agent, and work carrying a convention your teams agreed, such as a branch prefix or a label. Nothing infers authorship from writing style, because that guess is not defensible.
Why do you call the number a floor?
Because agentic work that comes from no recognised account and carries no agreed convention is invisible to any honest method. The reported figure is what can be proven, so the real rate is at least that and may be higher. We would rather tell you that than publish a number that cannot survive a challenge.
What is a disposition?
The verdict for a system, taken from where it sits on two axes: how much pressure there is to replace it, and how feasible replacement actually is. The four outcomes are Rebuild, Contain, AI-maintain and Maintain, and each one comes with the criteria that produced it.
Can we change the weights behind readiness?
Yes. Start from Balanced, Effort reduction, Modernisation or Risk reduction, then edit the individual weights directly. They rescale to a hundred on save and the ranking recalculates, so you can argue the priorities rather than accept ours.
Do you integrate with Jira?
Not directly. Every report and priority list exports to CSV and Excel in a shape that imports cleanly into Jira, so you keep one backlog rather than two.

See where agents should go next

Book a demo and see your own systems ranked by disposition, readiness and the weights that matter to you.

Book a Demo