Real implementation stories from the trenches. Classical QE practices evolved with PACTS principles.
No hype, no vendor speak—just what actually works in production.
Quality isn't tested in—it's built in. We're moving from testing-as-activity to agents-as-orchestrators,
bridging classical QE with agentic intelligence through PACTS principles.
🔨
Practitioner First
Everything here is battle-tested in production. Real implementations, actual failures, honest lessons.
No ivory tower theory—just what works (and what doesn't) from someone in the trenches.
🌍
Community Driven
Building the Agentics Foundation community in Serbia and sharing globally. Learning happens together—through
meetups, open source, and honest conversations about what quality means in the AI age.
Two months of parallel agents, a delivery pipeline I had to simplify, a QE organism finding its
way into daily work, and a live meetup where a couple of minutes of exploratory testing found two
bugs after an agent repaired 83 drifted tests.
Two weeks of hardening a platform at Ruv's pace, six podcast episodes, a first-ever Vienna
meetup, four releases about honesty, a website audit that turned on its own author, and the
fortnight a musician left the orchestra without the music stopping.
Two weeks of building quality nets around a platform racing toward release, a trip to Munich, a
benchmark where the frontier model never won a task, and the moment my own AI reviewers ruled
against me — QE-Court's first case was its own release, and the verdict was BLOCK.
Three weeks of beta-testing Reuven Cohen's MetaHarness, benchmarks proving cheap local models can
generate tests without losing quality, two meetups in two formats including our first panel, and the
week Adam and Klara came to Novi Sad.
Domain-Driven Design architecture with 60 specialized agents across 13 bounded contexts.
Queen Coordinator for hierarchical orchestration, TinyDancer 3-tier model routing, and ReasoningBank learning with Dream cycles.
TinyDancer 3-Tier Routing
Haiku/Sonnet/Opus intelligent routing with confidence-based escalation and cost optimization
Python reimplementation of the AQE framework using LionAGI orchestration. 18 specialized agents with
async-first architecture, alcall integration, and 82% test coverage.
LionAGI Native Integration
Builder pattern and Session management for persistent multi-agent coordination
99%+ Reliability
alcall integration with exponential backoff and fuzzy JSON parsing (95% error reduction)
Async-First Architecture
Real-time progress streaming and parallel execution with <1ms tracking overhead
v1.2.0 • Python 3.10+ • LionAGI Framework • MIT License
Sentinel: Multi-Agent Testing
Open-source agentic testing framework built with Rust and Python. Specialized agents working in concert—
functional testing, security injection, performance planning—all with explainability first.
PACTS-Based Architecture
Proactive, Autonomous, Collaborative, Targeted, Structured from the ground up
A Rust-native, local-first knowledge system that captures reusable patterns, retrieves them in context,
and updates their quality from real outcomes. It is the portfolio's meta-memory layer—not another test runner.
Living hypotheses
Patterns carry confidence, reward, reuse, and classified failure outcomes.
Hybrid retrieval
Full-text and neural search surface relevant prior learning.
Self-improvement
Decay, consolidation, surprise, and recommendations keep memory useful.
Local ownership
Rust, SQLite/SQLCipher, optional PostgreSQL, backup, and sync.
CaptureStore the problem, solution, context, domain, and provenance.
02
RetrieveBring relevant patterns back when a similar decision appears.
03
ObserveRecord success or a classified failure after the approach is used.
04
ImproveStrengthen useful patterns, expose weak ones, and consolidate carefully.
Reproducible findings and explicit limits
Experiment ledger
Published observations from Agentic QE, Nagual, MetaHarness evaluation, adversarial release work, and live exploratory sessions. Each result links to its context; none is a universal promise.
Judge qualification · controlled test
A 90% judge with zero recall was rejected
A synthetic judge scored 90% aggregate correctness while missing every evidence-corruption case. Qualification by fault slice rejects it: a healthy average cannot buy authority over the category it never catches. A test of the rule, not a field benchmark of a model.
83 test repairs, then two bugs in a couple of minutes
At a public meetup, a sub-agent repaired 83 tests broken by drift. A couple of minutes of manual exploratory testing then found two product bugs the suite had not established. One session on one app, recorded with its rough edges.
Across seven QE tasks and three routing policies, the cheap tier matched or beat the frontier tier on every task at 3–4× lower cost. The session cost $0.42; the sample is deliberately small.
On a 150-pattern Nagual set, qwen3:8b and gemma4:12b-mlx separated useful knowledge from lifecycle noise locally, with no API call. This is one curated dataset, not a general model ranking.
QE-Court's first case reviewed its own Agentic QE release and returned BLOCK after finding three data-loss or crash defects behind a passing unit suite.
Nagual crossed 300 accumulated patterns, then shifted the question from volume to lifecycle quality: retrieval, outcome evidence, decay, pruning, and whether learning survives real integration paths.
Growing the Agentics Foundation community in Serbia and beyond. Monthly meetups, open discussions,
and learning together about quality in the age of intelligent agents.
Agentics Foundation Global Meetups
Join the worldwide Agentic QE community. Monthly meetups happening across the globe—from Novi Sad to San Francisco,
building the future of quality engineering together.
Member of the global Agentics Foundation. Building chapters worldwide, bringing PACTS principles and agentic
engineering to quality practices across continents.
Looking for a speaker on Agentic QE, PACTS principles, or bridging classical to modern quality practices?
Let's talk about bringing practical insights to your conference or team.
Questions about Agentic QE? Want to discuss consulting or speaking opportunities?
Let's connect.
Stay Sharp in the Forge
Weekly insights on Agentic QE, implementation stories, and honest takes on quality in the AI age.
No spam, no vendor pitches—just practitioner-to-practitioner learning.