From Fragmented Tools to RACE Programming: What Process Automation Looks Like in a Real Software Project
Most software teams do not wake up one morning and decide to automate an entire delivery process. Automation usually starts somewhere less ambitious: a spreadsheet that is always out of date, a client asking for status updates, a Slack channel that has become impossible to follow, or engineers spending too much time translating vague feedback into actionable work.
That was the situation on one long-running enterprise software project.
Over roughly two years, the project evolved from three disconnected tools into a more integrated workflow in which client feedback, project context, work-item creation, delivery metrics, and early-stage AI-assisted engineering began to operate as parts of the same system.
The important part was not the technology itself, but the sequence.
Each layer of automation appeared because an existing process created enough friction to justify changing it.
The starting point: three tools, one human integration layer
The project relied on three main systems.
- Azure DevOps contained the development work: tasks, source code, CI/CD pipelines, and project documentation.
- A Google sheet gave the client a familiar way to follow planning.
- Slack was where the technical teams communicated.
Each tool worked reasonably well on its own, but they did not work together.
Updates made in Azure DevOps did not automatically appear in the spreadsheet. Conversations in Slack were not directly connected to work items. The client could not easily see what was planned for a release without entering a system designed primarily for the development team.
Someone had to keep the systems connected manually.
That meant copying information, reconciling statuses, answering repeated questions, and translating between technical language and the language used by business users.
It also created a measurement problem. The team had no simple shared view of what was entering the delivery process or whether reported issues were being resolved quickly enough. This challenge is common: the Stack Overflow Developer Survey 2024 found that non-coding work, including coordination, documentation, and status reporting, consumes a significant share of engineering time across organizations of all sizes.
The first useful automation therefore did not try to automate engineering.
It automated visibility.
Step one: make the release understandable
The first new layer was a release dashboard.
Its purpose was straightforward: give users one place to see what had shipped, what was in the current release, and what was coming next without requiring them to navigate Azure DevOps.
The dashboard showed a release timeline, searchable tasks, delivery dates, sprint information, and descriptions written for non-technical users.
That last point became particularly important.
The project effectively operated in two languages.
One was the language of development teams: bugs, implementation details, work-item structures, technical titles, and engineering terminology.
The other was plain English: what changed, why it mattered, and when users could expect it.
Rather than asking someone to rewrite every technical item manually, the team began using an LLM to translate project information into clearer descriptions. The model was supported by accumulated project context, including code, documentation, customer materials, and domain-specific information.
The technical system remained the source of truth. The AI layer changed how that information was presented.
For a user base of more than 300 people, that distinction mattered. Users did not need access to the internal engineering system to understand what was happening.
The useful failure: building the interface the client did not want
The next step looked sensible on paper.
The team built a simplified planning board connected directly to Azure DevOps. It showed work by person, supported filtering and priorities, and was intended to replace the client’s spreadsheet.
It did not.
The client continued using the spreadsheet.
That failure became one of the more practical lessons from the project: process automation does not automatically justify changing the interface people already trust.
Instead of forcing adoption, the team changed direction by connecting the spreadsheet to the underlying project data through Apps Script.
The spreadsheet stayed.
The manual maintenance did not.
This was a smaller change than replacing the tool, but it solved the actual problem.
The board remained useful during planning sessions, while the client could continue working in a familiar environment.
The system adapted to the workflow rather than asking the workflow to adapt to the system.
When Slack stops being a conversation and becomes an intake system
The next problem appeared as the application was opened to a much wider group of users.
The team expected more feedback, more questions, and more issue reports. A normal Slack channel was not designed to handle that volume cleanly.
Previously, reports arrived as free-form messages. Some were detailed. Others were not. Conversations became difficult to follow, and development teams often had to ask basic questions before they could begin investigating an issue.
The team introduced a structured Slack application for reporting problems.
Instead of writing an open-ended message, users submitted a small form containing the reporter, severity, description, and supporting files.
That changed the behavior of the system almost immediately.
Reports arrived in a consistent format. Status could be represented directly on the original message. A reported issue could move from new to in progress, to testing, to closed without forcing the reporter to search elsewhere for an update.
Slack was no longer only a conversation layer.
For this workflow, it had effectively become the front end of a lightweight ticketing system.
The benefit was not simply a cleaner presentation. Better structure at the point of intake improved the quality of the information reaching the delivery team.
According to the project team, some items previously required three or four days of clarification before meaningful work could begin. With the more structured process, even difficult cases generally required no more than several hours of clarification.
Project context became infrastructure
Structured intake solved only part of the problem.
A ticket still had to become engineering work.
That meant determining whether something was a bug or a change request, setting its severity and priority, assigning it correctly, writing a useful title and description, and attaching it to the right place in the project hierarchy.
This is where the project’s accumulated context became increasingly valuable.
Over two years, the team had built a context base containing project information, reference materials, internal documentation, domain terminology, and development history.
The model was no longer working with an isolated sentence from a user.
It was working with a representation of how the project itself operated.
That allowed the workflow to move in the opposite direction from the release dashboard.
Earlier, AI had translated technical information into plain language.
Now it could translate user language into structured business and engineering artifacts.
Through the project portal, a team member could review a report, classify it, generate a clearer description, determine priority, assign an engineer, and create the corresponding Azure DevOps item.
The system then posted the link back into the original Slack thread.
Everyone saw the same state without manually reconciling three systems.
This is one of the more transferable ideas from the project: the value of the AI layer increased as the context around the workflow became more structured.
The model itself was only one component.
The larger asset was the context that allowed it to act usefully.
This principle also sits behind First Line Software’s broader work in AI-Accelerated Engineering, where AI is embedded into engineering workflows while human engineers retain responsibility for architecture, quality, security, and production decisions.
Once the workflow is connected, it becomes measurable
Connecting intake and delivery created another possibility: measurement.
Before the automation, the question “Are we keeping up?” was partly subjective.
Once reported issues, work items, statuses, and delivery outcomes were connected, the team could measure both sides of the flow.
Intake metrics showed what kind of work was arriving, including the relationship between bugs and change requests.
Delivery metrics showed how much reported work had actually been closed.
That gave both the client and the delivery team a clearer signal when work started accumulating.
A dip in performance could be investigated rather than debated.
A shift in the nature of incoming work could also become visible. In this project, the team observed that change requests had begun to make up more of the incoming work than defect fixing.
The process had moved from anecdotal status reporting toward operational evidence.
The next layer: from creating one ticket to planning the work
The next stage of automation was still being developed when the project was presented.
Instead of creating a single item from a reported problem, the team began testing a wizard capable of proposing the work structure itself.
A report could be interpreted, given a title and description, and broken down into the likely implementation tasks.
The system could identify whether back-end or front-end work was required, create a QA task, and propose where the work belonged within the existing feature and user-story hierarchy.
This sounds like a minor administrative improvement until the volume is considered.
Small acts of project coordination accumulate.
Every correctly created task means less time spent navigating the tracker, recreating project structure, deciding where something belongs, or correcting work that was filed inconsistently.
More importantly, a structured work item becomes usable by the next stage of automation.
From automation around engineers to automation of routine engineering
The most experimental part of the project was a prototype internally referred to as a “Fixer”: an AI-assisted engineering component intended to handle selected routine changes.
The intended process contained three safeguards.
First: triage.
The system had to determine whether a reported item was appropriate for automated handling at all. If not, the work went to a human engineer.
Second: work with proof.
The intended pattern was to reproduce the problem through a failing test, make the change, and then run the relevant tests again.
Third: a release gate.
A human engineer remained responsible for deciding whether the generated change should proceed.
This was not presented as fully autonomous software delivery.
It was a controlled attempt to identify which parts of routine engineering could be delegated without removing human responsibility from the process.
That human-in-the-loop approach is also central to FLS’s AI Lab, which focuses on production-oriented agentic systems, governed execution, and human oversight.
The team also tested the concept against historical project work.
A set of 499 completed work items was used for comparison. The prototype’s output was evaluated against code that engineers had already implemented and deployed.
Without pre-filtering the tasks, roughly 43–45% were assessed as work the prototype could have completed automatically, a finding consistent with GitHub’s research showing AI assistance materially accelerating routine engineering tasks across different organizations.
The team did not treat that figure as a production-readiness threshold. The stated goal was to improve the rate substantially, strengthen automated test coverage, and keep human gating in place before introducing the approach into the live delivery flow.
What RACE Programming begins to look like in practice
The project described this direction as a series of steps toward RACE Programming.
RACE Programming is First Line Software’s AI-native software delivery framework. It shifts the focus from isolated coding assistance toward a delivery model built around structured context, executable work, automated validation, and human-controlled decision points.
The operating model starts with context.
Project knowledge is accumulated in a form that machines can use.
Then the task becomes executable.
A piece of feedback is not merely summarized; it can be transformed into a proposed set of tasks, attached to the appropriate project structure.
Then guardrails are added.
An automated engineering component can attempt suitable work, but only within defined boundaries, with tests and human approval.
Seen this way, the transformation is not about replacing an engineer with an agent.
It is about reducing the distance between a reported need and a verified implementation.
That distance normally contains many small handoffs:
- a user explains a problem;
- someone interprets it;
- someone creates a ticket;
- someone clarifies it;
- someone categorizes it;
- someone assigns it;
- an engineer investigates it;
- the work is tested;
- the client asks for an update;
- someone reports the status.
The project’s automation progressively shortened those seams.
A related example can be seen in FLS’s AI-Accelerated Engineering for Legacy Modernization case study, where AI-assisted engineering is also combined with structured validation and human-controlled delivery.
The broader lesson: automate friction, not people
The most useful pattern in this case did not come from starting with an AI strategy.
It came from repeatedly asking where information was being lost, duplicated, rewritten, or manually carried from one system to another. McKinsey’s research on AI-fueled software development identifies the same pattern: the organizations that extract the most value from AI in engineering are those that redesign the workflow first, rather than overlaying AI on top of an existing broken process.
The release dashboard addressed visibility.
The spreadsheet integration addressed adoption.
Structured Slack intake addressed ambiguity.
The context base addressed translation.
The work-item wizard addressed coordination.
The Fixer prototype began addressing routine implementation.
Each step built on the one before it.
That order matters.
Trying to automate implementation before the project has reliable intake, structured context, clear work-item relationships, test coverage, and release controls would simply move ambiguity further down the pipeline.
The experience suggests a more grounded way to think about AI-enabled software delivery.
The model is rarely the main constraint.
Rather, context, process design, integration, and governance are.
And the strongest automation often appears in the seam between a person and a system: where someone would otherwise have to translate, copy, classify, reconcile, or explain the same information again.
That is where a disconnected software process can begin to behave like a coherent system.
And it is also where AI-assisted engineering starts to become more than faster coding.
FAQ
What is RACE Programming?
RACE Programming is an AI-native software delivery approach that uses structured project context, executable work definitions, automated validation, and human-controlled decision points to reduce friction across the software development lifecycle.
How is AI-Accelerated Engineering different from AI code generation?
AI-Accelerated Engineering applies AI across the delivery process, not only during coding. It can support requirements interpretation, work-item creation, testing, documentation, workflow coordination, and implementation while keeping engineers responsible for architecture, quality, security, and release decisions.
What should software teams automate first?
The project suggests starting with repetitive handoffs and information gaps rather than coding itself. Release visibility, structured feedback intake, status synchronization, and work-item creation can establish the context and process discipline needed for deeper engineering automation.
Can AI automatically fix software issues?
Some routine work may be suitable for AI-assisted implementation, but the project used triage, automated tests, and human release gates. In historical testing across 499 work items, the prototype was assessed as capable of handling roughly 43–45% of unfiltered tasks, but the team did not consider that sufficient for ungated production use.




October 2026