Many mid-to-large enterprises are now discussing AI4SE: whether to standardize on a Coding Agent, build an enterprise Harness, adopt Spec-Driven Development, or bring Agentic Engineering into the delivery process.

The outcome, however, is rarely determined by the tool choice alone. It is determined by whether the organization can survive the early period in which confidence is most likely to disappear.

That period resembles the familiar DevOps transformation pattern. New practices may already be creating local improvements while overall delivery performance temporarily declines. Teams need to learn new ways of working, add specifications, tests, and reviews, and turn knowledge that used to live in individual experience into explicit assets. AI makes implementation faster, but it does not automatically remove the cost of decisions, verification, governance, and collaboration.

Organizations can therefore stop just as a new capability is beginning to form.

Managers see cycle time increase and conclude that the transformation has no value. Teams see more process and conclude that AI is making work slower. Procurement sees that tools have been deployed and concludes that the transformation is complete. A capability-building effort that needed consolidation is then labeled a failure, or prematurely packaged as a success.

This is the J-curve problem of AI4SE transformation.

Before asking how to survive that dip, there is a prior judgment: whether the organization has placed AI4SE in the right problem domain. If that cognitive framing is wrong, even a well-structured path will be executed with the wrong response strategy—and failure is often only a matter of time.

Drawing on our consulting experience in mid-to-large R&D organizations and the experience report From AI4SE Pilot to Conditional Scale: A Transformation-Path Experience in Mid-to-Large R&D Organizations, this article presents a more durable path:

Explore first, to see whether the method can run; Consolidate next, to turn one successful run into a capability the team can repeat and teach; then Scale in waves, so new teams can operate independently within guardrails.

The crucial point is that these stages should not advance by calendar alone. Evidence should decide whether the organization is ready for the next stage—and whether that evidence can be read correctly depends on first classifying the problem domain correctly.

1. Understanding the J-Curve

Early DevOps transformations often include a temporary decline. When teams introduce continuous integration, automated testing, infrastructure as code, visualized flow, and cross-functional collaboration, hidden problems become visible:

  • waiting, rework, and manual operations are exposed;
  • teams need time to learn new tools and coordination habits;
  • quality controls that depended on individual experience need to become repeatable mechanisms;
  • shorter feedback cycles require investment in tests, pipelines, and environments;
  • local changes temporarily increase coordination costs across adjacent teams.

The short-term picture gets worse before the system has a chance to move into a higher-performance state.

AI4SE amplifies this pattern. A Coding Agent can quickly generate code, tests, and documentation, but it does not automatically answer:

  • What problem are we solving?
  • Which constraints must not be violated?
  • Does the generated result match the real business intent?
  • Who owns acceptance and risk decisions?
  • How will this learning be preserved for the next team?

It is useful to think in terms of two clocks:

Code-production clock: increasingly fast
Definition, decision, verification, and governance clock: not automatically faster

If management watches only code-generation speed, the organization gets a confusing result: more code, but no faster delivery; stronger tools, but more review and testing work. The bottleneck has moved from implementation to definition and verification.

An early decline therefore does not necessarily mean that the direction is wrong. It may mean that the organization is paying the learning, quality, and governance costs of a new method. The real question is whether those costs are being converted into reusable organizational assets, or merely into more meetings, documents, and manual work.

2. First Classify the Problem: Cynefin as a Cognitive Lens

Many transformations fail not because execution is weak, but because the problem is misclassified at the beginning. An organization treats a complex problem as if it were simple, or continues experimenting after the problem has become stable enough to standardize. Once the response strategy is mismatched to the problem domain, failure is often only a matter of time.

The Cynefin Framework is useful here as a cognitive lens. It does not tell management which AI4SE tool to buy. It helps answer a prior question: Are we facing a problem for which a known best practice can be applied, or one in which the method must be discovered through bounded experimentation?

The following simplified mapping is enough for transformation design:

Cynefin domainProblem characteristicSuitable responseTypical AI4SE meaning
ClearCausality is stable and a best practice is knownSense - Categorize - RespondStable standard workflows, low-risk automation, and explicit gates
ComplicatedCausality exists but requires expert analysisSense - Analyze - RespondA Playbook exists but needs adaptation to local technology and organizational constraints
ComplexCausality becomes clear only in retrospect; methods emerge through practiceProbe - Sense - RespondEarly AI4SE pilots exploring process, methods, human-Agent roles, and Harness
ChaoticThe system has lost stability; order must be restored firstAct - Sense - RespondMajor quality, security, or delivery instability requiring containment before transformation experiments

Cynefin also warns about a fifth condition: disorder, when the organization has not yet determined which domain it is in. The first action is then to decompose the situation and distinguish what is clear, what needs expert analysis, and what still requires exploration.

Why early AI4SE is usually complex

In mid-to-large enterprises, early AI4SE work is often complex rather than clear:

  • processes, systems, technology stacks, and quality constraints differ across organizations;
  • model, tool, and Harness capabilities are still evolving;
  • human-Agent accountability does not appear automatically after a tool is purchased;
  • requirements, verification, permissions, and knowledge capture influence one another;
  • a method that works for one team may not transfer directly to another.

The pilot should therefore not pretend that a universally reusable best practice already exists. It should operate as a bounded probe: choose a real but controlled scenario, run an end-to-end method chain, observe what works and what fails, identify the conditions that matter, and turn the findings into inspectable assets.

This is also why “buy one tool, train everyone, and require adoption” is usually a poor starting point. Those actions disguise a complex problem as a clear one. They assume that the answer already exists and that the organization only needs to execute it. The unresolved questions about definition, verification, and accountability then surface later, at much greater scale.

The three stages as a domain transition

The Cynefin lens gives another interpretation of Explore, Consolidation, and Scale:

Explore: probe - sense - respond in the Complex domain
         discover a locally feasible method through real work

Consolidation: move stable findings toward the Complicated domain
               use Playbooks, Harness hardening, and coaches to repeat and adapt

Scale: replicate only what is stable enough to transfer
       diffuse in waves within explicit guardrails and feedback loops

Two opposite mistakes are equally dangerous.

The first is treating a complex problem as clear. Management announces the “right tool,” “right process,” and “right training package,” then measures adoption. The organization amplifies unverified assumptions before it has understood the problem.

The second is keeping every problem in the Complex domain forever. Teams keep experimenting, switching tools, and rewriting prompts, but never turn stable practice into Playbooks, gates, roles, and Starter Kits. Exploration becomes a refusal to standardize, and no transferable capability forms.

The right transformation is neither immediate standardization nor permanent experimentation. It allows the problem domain to move as evidence accumulates: explore first, consolidate next, and Scale only what has earned transferability.

3. Three Common Failure Paths

These three paths look like execution problems. In Cynefin terms, they are usually domain mismatches: treating a Complex-domain problem as if it were Clear, or scaling with Clear-domain tactics before the domain transition has been earned.

Tool-only transformation

The most common approach is to select a tool, purchase it centrally, install it everywhere, and require every team to use it.

In Cynefin terms, this disguises early AI4SE as a Clear-domain problem: causality is assumed to be stable, a best practice is assumed to exist, and the organization only needs to categorize and execute. The implicit assumption is simple:

If everyone uses the same tool, the organization will naturally develop the same capability.

Tools solve the problem of access. They do not solve the problem of how the capability should be embedded in work.

The organization still needs to redesign:

  • how requirements become AI-executable specifications;
  • which context, tools, and permissions an Agent can access;
  • how generation, application, and verification are separated;
  • how code, tests, specifications, and decisions remain traceable;
  • how security, compliance, and audit cover generated artifacts;
  • how effective practice becomes rules, skills, templates, and runbooks.

Without these decisions, the organization gets faster local experimentation rather than AI4SE. Each team creates its own prompts, context packs, and checking habits. The result is higher organizational entropy: several ways to do the same thing, quality that depends on a few experienced individuals, and knowledge that leaves with key people.

Tool procurement can be an input to transformation. It cannot be treated as the transformation itself.

Training-only transformation

The second approach is to run a rapid training program, teach everyone the tool and a set of prompting techniques, and declare the capability established when the course is complete.

This is another Clear-domain mismatch: it assumes that the answer is already known and that the remaining work is to transmit it to everyone. Training matters, but course completion only proves that people attended. It does not prove that the organization can operate a new way of working.

AI4SE is a work discipline embedded in real delivery. Teams need to learn through actual requirements:

  • How precise must the specification be?
  • Which tasks belong to the Agent and which must remain human-led?
  • Who verifies the Agent’s output?
  • What happens when verification fails?
  • Which learning should be written back to the Living Spec, rules, and Harness?

Training without real requirements, codebases, and quality constraints easily becomes a tool demonstration week. Participants leave with useful impressions but no clear first change to make in the delivery system.

Training should therefore be a supporting action within transformation, not its acceptance criterion. A more effective pattern is for external coaches to guide teams through real work, introduce the minimum knowledge needed at the moment of need, and preserve decisions as reusable assets.

Pilot-equals-scale

The third failure mode is the most subtle:

The pilot finished on schedule and received positive feedback, so the organization is ready for enterprise rollout.

This skips the Complex → Complicated domain transition: a successful local probe does not yet mean the organization has a method stable enough to analyze, adapt, and transfer. A pilot usually answers:

Can this method run under the current team, project, coach, and context?

Scale must answer a harder question:

Can another team reproduce the method in its own environment with less dependence on the original experts?

These are different questions.

Between pilot completion and Scale readiness lie method assetization, toolchain hardening, explicit ownership, internal coach development, sponsor cadence, and new-team validation. Skipping that work means handing the first team’s tacit knowledge to the second team and expecting the second team to fill every gap locally.

The first team becomes a transformation expert group. Other teams become dependent on it. When the external coach leaves, the capability leaves as well.

4. From “Can Run Once” to “Can Be Reproduced”

The paper separates two types of uncertainty.

Product uncertainty asks whether AI4SE can improve the selected work under local business, codebase, and quality constraints. Real pilots reduce this uncertainty.

Transfer uncertainty asks whether another team can reproduce the method without reinventing it. Consolidation and new-team validation reduce this uncertainty.

The distinction explains why “the pilot succeeded, but Scale admission is not yet granted” is a coherent outcome. The two decisions answer different questions.

Management should also distinguish three readiness states:

StateMeaningWhat it supports
CitableThe method is stable enough to show, discuss, and trace, with gaps made visibleExplore closeout and entry into Consolidation
RunnableThe main path works reliably in the intended local workflowEvaluation of whether the method can be handed to another team
TransferableA new team can complete a cycle with declared guardrails and bounded coachingAdmission of a Scale wave

A directory full of commands, skills, Agents, and hooks is not proof that the assets are runnable. A path that works with a consultant beside the team is not proof that it is transferable.

5. AI4SE Reorganizes the R&D Operating System

In practice, we observe AI4SE through four connected directions:

DirectionQuestionTypical concerns
Spec-Driven DevelopmentHow is intent made precise?Living Spec, Change Request, Delta Spec, acceptance scenarios
Agentic EngineeringHow do humans and Agents complete work together?Task orchestration, context use, feedback, escalation
Harness EngineeringHow does the Agent work reliably inside boundaries?Tools, permissions, knowledge, feedback, observability, gates
Operating ModelHow does local practice become organizational capability?Roles, coaches, assets, measurement, rollout

These are not four independent toolkits.

They also need to be understood through two cross-cutting concerns.

The first is Effectiveness. It does not mean speed alone; it includes quality, trust, security, compliance, and valuable outcomes after waste is removed. The second is Harmony, the human-Agent collaboration contract: who owns intent, who maintains the specification, who orchestrates tools, who verifies the result, which risks can be handled under supervision, and which risks require direct human participation.

AI4SE is therefore not a new “AI layer” stacked on top of traditional engineering. Process, methods, and tools still have distinct responsibilities, but the executors are now a combination of humans and Agents. Effectiveness is the underlying outcome, while Harmony cuts across the entire delivery path.

SDD as a verifiable change contract

When AI participates heavily in implementation, a work item cannot remain only a sentence in a backlog or a loosely understood user story. Intent, scope, constraints, acceptance conditions, and risks need to become a sufficiently precise change contract.

One representative flow is:

Business intent
  -> Change Proposal
  -> Delta Spec
  -> Design and Task Plan
  -> Generation and implementation
  -> Independent verification
  -> Verified Increment
  -> Merge back into the Living Spec

The value of SDD is not more documentation for its own sake. It gives humans and Agents a shared source of truth. Code is an implementation snapshot of the specification, not the only source of truth.

Agentic Engineering as governed participation

Agentic Engineering is not simply connecting a chatbot to an IDE. It lets Agents participate in repository exploration, solution design, task breakdown, test completion, build repair, review assistance, and knowledge capture.

Participation does not mean independent accountability. Human and Agent responsibilities need to be explicit:

  • the Intent Owner defines goals, boundaries, and risk acceptance;
  • the Spec Steward turns intent into a verifiable specification;
  • the AI Orchestrator manages context, tools, and execution paths;
  • the Verification Lead owns independent verification and the final quality judgment.

One person may combine these functions, but none may be left vacant. In particular, the Agent that applies a change must not be the sole authority that declares the change correct.

Harness Engineering as a controlled environment

A Harness is not a collection of prompts or a plugin directory. It is the external control layer organized around the delivery method:

  • how context and knowledge are supplied;
  • which tools an Agent may call;
  • how permissions and data boundaries are controlled;
  • how the process is observed and fed back;
  • which quality gates are mandatory;
  • how learning becomes rules, skills, templates, and runbooks.

Harness design should therefore be driven by method and workflow. It should not begin with a platform purchase followed by an attempt to force every process into it.

SDD’s Align, Verify, and Merge reviews control an individual change. Explore closeout and Scale admission are organizational stage decisions. These two control loops must not be conflated.

SDD also does not require an organization to discard Scrum, Kanban, or another existing team operating model. Agile and Lean still provide principles for flow, feedback, and continuous improvement. Scrum, Kanban, or Scrumban can still provide team cadence and collaboration structure. The main shift is in the work-item contract: from “there is a story card” toward “there is a traceable and verifiable change contract.”

6. Explore: Prove That the Method Can Run

The purpose of Explore is not to prove that every team should use one identical solution. It is to explore a feasible method chain within a controlled scope.

Explore should examine at least four dimensions:

  1. Process: how requirements, design, implementation, verification, delivery, and knowledge capture connect;
  2. Tools: how Agents, MCP, CLI, CI, tests, and permissions work together;
  3. Methods: how SDD, Harness Engineering, and Agentic Engineering enter real work;
  4. Human-Agent division of labor: what humans decide, what Agents execute, and what must be independently verified.

The pilot should use a real but bounded Change Request as its validation vehicle. The question is not whether every business feature ships. The question is whether the chain from intent to specification, execution, verification, and learning can be exercised.

At Explore closeout, management should ask:

  • Is there a citable initial Playbook?
  • Can opportunity, specification, and best-practice assets be inventoried?
  • Are executable assets honestly labeled Ready or Stub?
  • Are effectiveness and maturity signals reported with their limits?
  • Is the Consolidation path concrete enough to discuss?

Explore may close with unfinished assets. It may not close with hidden gaps.

In the paper’s Org-A case, Explore took approximately six weeks and covered diagnosis, scenario selection, four workshop weeks, twelve opportunity themes, and a real change request. Management accepted Explore closeout because the method chain had run end to end under real constraints and the initial method and opportunity assets could be cited. The main executable chain was still largely stubbed, however, and internal coach and cross-team transfer evidence had not yet been established.

The correct conclusion was therefore “enter Consolidation,” not “roll out everywhere.”

Adaptation matters in complex process environments

The Org-B material in the paper provides another important warning. In a manufacturing organization with incumbent Integrated Product Development (IPD) and SAFe-class coordination constraints, AI4SE cannot simply demand that the organization abandon its existing process and adopt a new AI process.

Org-B is documented primarily as a late-Explore adaptation design, not as completed organization-level validation. The work focused on connecting existing requirement entry points and IPD work items to SDD change contracts at the team level. The paper does not present it as formal Explore closeout, completed Consolidation, or Scale admission.

Scale therefore needs to preserve two things at once: the core admission criteria must not be diluted, but process entry points, role arrangements, and Harness assets must be adaptable to local conditions. What transfers is not an immutable process manual. It is a set of inspectable principles, artifacts, responsibilities, and admission conditions.

7. Consolidation: Turn One Success into a Teachable Capability

Consolidation is the stage most likely to be compressed, and the stage that most determines whether Scale will work.

It is not another demonstration. It is repeated use in real iterations, turning failures and friction into better methods and assets.

Consolidation should:

  • repeat the method on real requirements, defects, and technical debt;
  • revise clauses that read correctly but do not run in practice;
  • harden the primary executable path from Stub to Ready;
  • make role ownership and verification responsibility explicit;
  • develop internal coaches;
  • preserve the separation between applying a change and verifying it.

The core exit condition is:

The team can not only use the method, but teach it to others.

That requires a more stable Playbook, a runnable main path, named ownership, team members who can explain the reasoning behind the method, bounded support for new members, and a gradual shift from external coaching to observation and correction.

In the Org-A Consolidation observation window, the team recorded 15 requirements, 18 changes, and 8 finished changes over five weeks. All 18 changes entered through the mandatory SDD path and produced specifications. These are process signals, not ROI evidence, but they show repeated use of the method on real demand.

The same evidence also preserved the remaining gaps: Verify did not yet have a complete set of Reject or Conditional records, Align took longer than expected, and one AI-orchestration role was still being hardened. Those are not failure statistics. They are evidence that Consolidation was not complete and that Scale should remain conditional.

8. Scale: Conditional, Wave-Based Diffusion

Scale should not mean moving every team at once. A safer pattern is wave-based diffusion:

Pilot team
  -> Gold seed
  -> First seed coaches
  -> New-team validation
  -> Next seed-coach wave
  -> Wider diffusion

Scale admission should require:

  • observed “team can teach” evidence;
  • a runnable main execution path rather than a mostly stubbed path;
  • a Starter Kit for receiving teams;
  • named ownership for intent, specification, orchestration, and verification;
  • sponsor cadence, time, and resources;
  • a new team that can complete a full cycle with bounded coaching.

Scale admission should also be reversible. If a new team cannot operate independently, the organization should return the deficient asset, role arrangement, or enablement mechanism to Consolidation instead of blaming the team or forcing the rollout forward.

The most important thing to grow during Scale is not the number of tool licenses. It is the number of capable seed coaches, the quality of method assets, and the number of new teams that can operate independently.

External coaches, gold seeds, and seed coaches

The role of an external coach is not to remain the permanent operator of transformation. It is to help the organization build its own coaching ladder:

  1. External coaches guide early pilots and help teams identify method, tool, and organizational constraints;
  2. Gold seeds emerge from Explore and Consolidation because they show strong learning ability, influence, and end-to-end understanding;
  3. Seed coaches learn to help other teams transfer and apply the method.

A gold seed is not merely the person who uses the tool most fluently. More important is the ability to explain why the method works, recognize its limits, and make tradeoffs among business intent, engineering quality, and human-Agent collaboration.

9. Maturity and Measurement as Evidence

A maturity model should not rank teams or require all teams to reach the same level at the same time. Its useful role is to reveal capability profiles and bottlenecks:

  • effectiveness goals and value hypotheses;
  • human-Agent workflow;
  • specification and verification discipline;
  • Agent runtime, context, and permissions;
  • security, quality, and accountability boundaries;
  • internal coaches, knowledge assets, and continuous improvement.

Assess stable behavior, not plans, demos, or one-off successes. One person doing something does not prove that a team can do it. One team doing it does not prove that another team can transfer it.

Useful measurement dimensions include:

DimensionExamples
Flow and efficiencyTime from requirement to delivery, PR cycle time, build repair time, waiting time
Quality and trustRework, defects, test and verification evidence, review issue density
Organizational capabilityAsset reuse, seed-coach capacity, new-team independence
GovernancePermission blocks, compliance issues, cost visibility, human accountability records

Code volume, Agent invocation counts, prompt counts, license counts, and training attendance may show activity. They do not prove capability.

Better management questions are:

  • Is the method being used in real work?
  • Can its artifacts be cited and traced?
  • Is the main path actually runnable?
  • Is verification authority still human-owned?
  • Can a new team complete a cycle without depending on the original team?
  • Can the organization improve the Playbook and Harness when the method meets friction?

10. A Management Checklist

Ready to close Explore and enter Consolidation

  • The method chain has run end to end in a real scenario.
  • A citable initial Playbook exists.
  • Opportunity and specification assets can be inventoried.
  • Ready and Stub states are visible.
  • Known gaps and the next-stage plan are explicit.
  • Local signals have not been packaged as organization-wide ROI.

Not ready for Scale

  • The main path still depends on a consultant operating it live.
  • The tool catalog looks complete, but the execution path is mostly Stub.
  • Pilot members know what to do but cannot explain why.
  • Internal coaches do not yet exist.
  • The person applying the change is also the sole verifier.
  • No new team has independently completed a cycle.
  • Management has budgeted for rollout but not for Consolidation.

Ready to admit a Scale wave

  • The pilot team can repeat and teach the method.
  • The Playbook, Starter Kit, and Harness main path are stable.
  • Roles and verification responsibilities are explicit for the receiving team.
  • A new team can complete a full cycle with bounded support.
  • Sponsors have committed resources, cadence, and retrospectives.
  • The organization can return failed assets or role arrangements to Consolidation.

Conclusion

The hardest part of AI4SE transformation is not helping a skilled team use an Agent in a successful demonstration. It is producing trustworthy software consistently while people, projects, and process constraints change.

That requires a different definition of “transformation complete”:

  • not tools deployed, but methods validated;
  • not training completed, but real work repeated;
  • not a successful pilot, but an independent new team;
  • not a short-term efficiency spike, but simultaneous improvement in quality, trust, flow, and organizational capability.

Explore, Consolidation, and Scale are therefore not merely three calendar periods. They address three different uncertainties:

Explore: Can the method run locally?
Consolidation: Can the team repeat and teach it?
Scale: Can another team run it independently while preserving quality?

Management’s most important responsibility is to give the method enough room to learn through the bottom of the J-curve while continually classifying the problem domain: what still needs exploration, what is ready for consolidation, and what has earned conditional Scale. Do not abandon the transformation because of an early dip. Do not launch an enterprise rollout because of a local success. And do not apply Clear-domain tactics to a Complex-domain problem before the domain has even been judged correctly.

What deserves to Scale is never just a tool. It is a way of working that can be verified, taught, transferred, and continuously improved.

Further Reading