My GitHub was on fire in August, even while I was on holiday. Multiple coding sessions running in parallel, agents building and checking different parts of the work, me trying to keep the direction coherent. Last year, I could not have imagined working at this tempo. This year, it became a practical engineering problem.
At Cognitum, each of us runs several agent sessions at once. The volume of change that a small team can produce is extraordinary. It also exposed every part of the organization designed for a slower rhythm, including the quality machinery I was helping build.
My agents had overengineered our CI/CD. Deploying new versions had become slower than it needed to be. I orchestrated that work, so fixing it was my responsibility too.
That is where most of this story lives: in the distance between being able to produce more software and being able to look after it properly.
Making the Way to Production Shorter
In the last article, I described the beginnings of a QE organism for Cognitum. The deterministic foundation was running; much of what I wanted above it was still being built and tested. August and September put that ambition against the less glamorous details of everyday delivery.
One of those details was repetition. Routine changes were going through layers of promotion and staging ceremony. Each layer could be explained individually. Together, they slowed the whole team down.
Agents are very capable of adding structure. Give them a problem involving risk, and they can produce another workflow, another approval boundary, another document explaining how the first two connect. I had to bring the work back to the question a quality engineer should ask of any control: what risk does this actually reduce, and what evidence does it add?
Some of the answer was subtraction. I worked on a more direct route from reviewed change to production in August, and continued making it leaner through September. We retained production checks and the ability to revert. We also gave more attention to what happened after deployment: monitoring, observability, and alerts that reached the people who could respond.
I do not have a measured deployment-speed improvement to put in a headline. What changed was the shape of the work. Less repeated ceremony around ordinary changes, more attention to verifying the running result and knowing what to do when it was wrong.
For a team generating changes in parallel, that matters. If quality work accumulates in a queue nobody can move through, people eventually spend their energy working around it. I want the quality system to help us keep pace without losing sight of what we are shipping.
The Work After the Green Check
The correction came with its own reminder. Having a rollback target available did not prove that the previous version could actually serve when we needed it.
It sounds obvious written that way. It is much easier to miss when the release machinery has successfully recorded a previous target and all the expected pieces appear to be in place. We needed to verify the runnable fallback, not stop at its existence.
That is a very old testing problem in a very current setting. We had checked a representation of readiness and still needed to check the behavior that mattered.
Alerting needed the same treatment. Can a problem be detected? Does the notification reach its destination? Can we observe recovery as well as failure? We exercised those paths, including delivery to Slack and email, rather than treating a configured channel as proof. The verification was specific to the paths we tested; it does not establish coverage of every possible production failure.
This is the kind of work that becomes more important as agents make implementation faster. A green check closes one question. Questions remain about the deployed version, the behavior a user encounters, and the response when that behavior changes. Those questions need owners and evidence too.
The Organism Learns to Listen
Ruv has created most of the supporting structure around this work: the Cognitum Slack environment, swarm orchestration, and federation foundations. My work is to build the QE harness into that environment, so quality plays a useful role in how people and agents already work together.
Part of the work has been building specialized skills: making the quality questions, expected evidence, and escalation boundaries explicit enough that a new agent session doesn’t have to rediscover the whole practice. Another part has been on-call support, helping with routine investigation and review while escalating decisions that need a human.
By late September, the organism was producing triage work that maintainers could correct, and those corrections were feeding improvements to the source information and context. That is useful progress. I am still withholding the larger claim that it reliably makes useful decisions. Automatic evolution needs independent evaluation and human authorization.
The connection with Core Memory is being built and hardened for the same reason. Observations and reviewed lessons should be available to the next session, with enough history to understand where they came from. Otherwise, parallel sessions can become very efficient at repeating one another’s mistakes.
But retaining an incident does not show that we learned from it. Retrieving it does not show that the next decision improved. That last step is where the quality question remains open, and it is where I want the evidence to lead us.
Readers who remember the orchestra will recognize the problem. Many more musicians are available now. The work around them still determines whether they can play together, and whether anyone notices when the performance goes wrong.
The Fleet Had to Face the Same Questions
The public Agentic QE project kept moving too: fifteen releases since the last article, from v3.13.5 through v3.14.6. The community did a substantial part of that work, especially in late September, when contributors kept finding places where the tools didn’t live up to their promises.
One example in v3.14.4 was painfully relevant to the way we now work. Parallel workflow steps could overwrite one another’s outputs. If several agents finish their work and the orchestration loses part of it, adding more agents will not help. The fix belonged in the shared machinery meant to preserve their results.
Another example came from the judge-qualification work. A controlled test constructed a judge with 90 percent aggregate correctness and zero recall on evidence corruption. The qualification rule had to reject it. A reviewer that misses every corrupted-evidence case cannot earn authority simply by doing well on the easier remainder.
That was a synthetic test of the rule, not a field benchmark of a model. Its value was in making a dangerous blind spot explicit. Experienced testers will recognize the reasoning immediately: the average can look healthy while the failure category you most need to catch remains completely unprotected.
Rudy (@rudycelekli) contributed much of the repair work across the late-September releases. Chris (@pacphi) brought reproducible reports, and Nick (@nagoodman) helped with retesting. That work changes the project in ways my own sessions cannot. Other people bring different environments, different assumptions, and patience for problems I have stopped seeing because I know the intended route too well.
I can orchestrate more coding than ever. I still need people who use the result and tell me where it fails them.
Two Bugs After Eighty-Three Repairs
Our fifteenth Serbian Agentics Foundation meetup on September 10 offered the most compact version of this lesson.
During the session, a Sonnet sub-agent fixed 83 failing tests caused by test drift. That is substantial maintenance work, and it is exactly the kind of work I am happy to have help with. To be clear about the number: those were 83 drifted tests, not 83 application bugs discovered and repaired.
Then I spent a couple of minutes doing manual exploratory testing and found two bugs. A comment count stayed at zero. A notification’s read state did not clear as it should. We fixed the comment-count issue, rebuilt, and verified it. The recording is public, including rough edges.
That sequence is worth showing to anyone trying to understand where testing is going. The agent handled a large amount of test maintenance. I went into the product and asked different questions by using it. The test suite and the exploratory session looked at different aspects of the same software, and the second activity still found something the first hadn’t established.
I keep returning to this format for the community because people can see the work change direction. We can talk about an approach, use it, encounter a problem, and investigate it together. That sequence offers more to learn than a demonstration where every step arrives pre-polished.
It also brought me back to a point I discussed after the meetup: as implementation and verification get faster, product discovery and stakeholder alignment deserve more attention. We can now build a misunderstanding at impressive speed. Spending time on what people actually need, and checking that understanding with them, has become an even more consequential part of the job.
What I Am Taking Into October
My GitHub activity graph tells one part of the story. The part I would rather carry forward is the correction: I had to simplify quality machinery that was getting in the way, then strengthen our ability to observe and respond to the software we delivered.
The QE organism is further along than it was in early August. The skills, on-call work, and Core Memory connection are becoming parts of a working practice. Some questions remain open, especially whether retained experience consistently improves later decisions. I would rather keep that question visible than give the system a more impressive description than it has earned.
For quality engineers, there is a great deal of familiar work here. Choose the risks worth examining. Check the oracle. Follow a result into the running product. Test the recovery you expect to depend on. Make it possible for a colleague, a contributor, or the next agent session to understand what actually happened.
I am building and testing at a pace I could not have imagined last year. These two months have made me more careful about what I ask that speed to accomplish.
Stay curious. Keep learning. Keep sharing. Knowledge is power.
This is the thirty-fifth article in The Quality Forge series. Previous: “The Orchestra Keeps Playing” described moving half the parallel sessions to Codex and the QE organism master plan. This one covers the two months after: simplifying an overengineered delivery pipeline at Cognitum, verifying rollback and alerting paths rather than their configuration, the QE organism’s skills, on-call support and Core Memory connection, Agentic QE releases v3.13.5 through v3.14.6, and the fifteenth Serbian Agentics Foundation meetup (recording). Cognitum is at cognitum.one.
Dragan Spiridonov is the Head of Agentic Quality Engineering @ Cognitum One, Founder of Quantum Quality Engineering, Secretary of the Agentics Foundation BoD, Chair of the Agentics Foundation Training Committee, and one of the AI Chapter leads for the Ministry of Testing. He is currently building the Serbian Agentic Foundation Chapter in partnership with StartIt centers across Serbia.