Internal tool · Nov 2025 — present
Aegis-QA — the tooling I built
An internal platform I built to manage the QA work itself: triggering runs, API testing, Playwright scripts, SQL queries, Datadog lookups and test-case management with automation-script recording, all in one place.
- Role
- Creator & maintainer
- When
- Nov 2025 — present
The problem
The work was spread across a dozen tools. Postman here, a Playwright repo there, SQL in a client, Datadog in a browser tab, test cases in a spreadsheet. Every investigation meant stitching them back together by hand.
What I did
- One place to trigger and watch runs across UI, API, flow and security layers.
- A recorder that turns a real session into a reviewable script instead of hand-written selectors.
- API tests, SQL queries and Datadog lookups available next to the run that needs them.
- Test-case management tied to the automation that covers it.
- Self-healing locators. Proposed and sandbox-verified, never applied without a human saying yes.

Why it exists
QA is my profession. AI is my curiosity. Aegis-QA is what happened when the two met: I did not wait for a vendor to sell me a tool, I learned enough AI to build my own.
That is the short version of how a QA engineer ends up maintaining a platform. The QA judgement is mine. AI is how one person shipped something that would normally need a team, and learning enough of it to do that is the part I would do again.
Try it
Open the full portal walkthrough ↗
A working demonstration, running in your browser. It talks to nothing: the run is on a timer, the failure is scripted, and every value is invented. The real tool drives live payment systems, and no capture of that belongs on a public page.
Run a suite
5 passed · 1 failed · mode suite
- checkout.guest_card4.1s
- checkout.saved_card2.2s
- refund.partial_amount2.8s
- payout.idempotency_key9.4s
- ledger.reconcile_on_read1.2s
- webhook.replay_is_ignored0.9s
Record a test
4 of 4 captured
What the recorder saw
- click · button "Add to basket"
- fill · field "Email" · acme@example.test
- click · button "Pay now"
- expect · text "Order ord_4821 confirmed"
What it compiled to
page.get_by_role("button", name="Add to basket").click()
page.get_by_label("Email").fill(email)
page.get_by_role("button", name="Pay now").click()
expect(page.get_by_text(f"Order {order_id} confirmed")).to_be_visible()Triage a failure
✗ payout.idempotency_key
stops here until a human approves
- ClassifiedLocator no longer matches. This is not a product defect.
- DiagnosedThe Continue button moved inside a dialog in build 218.
- Fix proposedget_by_role("button", name="Continue") scoped to the dialog.
- Sandbox verifiedRuns green in isolation. That is not proof the bug is gone.
- Waiting for a humanNothing is applied until someone approves it.
What it replaced
Before this, the job was five windows. The comparison worth making is not against a product with a pricing page. It is against the stack a QA team actually runs when nobody has built them one.
- 01
Write a test
Before
Hand-written selectors, rewritten every time the UI moved.
With Aegis
Record the real session. The spec is generated, and regenerable from the recording.
- 02
Run it
Before
A framework repo driven from a terminal, results scrolling past in the console.
With Aegis
One place, with the run history and its artefacts kept together.
- 03
Check the API
Before
Postman in another window, with its own environment files to keep in sync.
With Aegis
Collections run beside the UI suite, against the same run record.
- 04
Confirm the data really changed
Before
A SQL client, queried by hand, when someone remembered to.
With Aegis
Saved queries against the state the run just produced. This is where the UI-says-success-but-the-ledger-disagrees bug gets caught.
- 05
A locator stops matching
Before
Find it, guess a new one, hope it holds until next week.
With Aegis
A proposed replacement, verified in a sandbox, waiting for a human to approve it.
- 06
A run fails at 2am
Before
Read the log, then the trace, then the diff, then ask the developer.
With Aegis
Classified and diagnosed on arrival, with the trace for the failed request pulled in beside it.
- 07
Know what is covered
Before
A spreadsheet, out of date the day after it was written.
With Aegis
Cases tied to the automation that covers them, so a gap is visible rather than assumed.
- 08
Test the Android app
Before
A separate project, on a separate framework, that nobody maintained.
With Aegis
The same record-and-compile path, over Appium.
What is in it
Session recorder
Records a real session on the real application and keeps the intent, meaning which element, which role and which label, rather than the raw selector that happened to match. The recording is the source of truth; everything downstream is generated from it and can be regenerated.
Compiler
Turns a recording into readable Gherkin and a runnable Playwright spec. Nav clicks, entity IDs and CRUD verbs are recovered from the session, so a generated test reads like one a person wrote.
Execution engine
Runs UI, API, flow and security layers across the five runner modes, headed or headless, locally or on the deployed box, with the run history and its artefacts kept together.
Self-healing locators
When a locator rots, the run proposes a replacement and verifies it in a sandbox before anyone sees it. A sandbox pass is advisory. It never resolves the failure, and never writes itself into the knowledge base on its own. A human approves, or it stays a suggestion.
Failure intelligence
Every failed run is classified, diagnosed, given a fix plan, and where the fix is safe and on an allowlist, executed and verified. Deterministic checks run first and AI runs last, so the cheap and certain answer is never skipped in favour of a guess.
Knowledge graph
A live model of the application built from what has actually been observed running, not from a diagram. It answers which tests cover a change, what a change puts at risk, and which surfaces nothing has ever touched.
API and database verification
Collections and contract checks run beside the UI suite, and saved SQL queries confirm the state a run just produced. This is where the UI-says-success-but-the-ledger-disagrees class of bug gets caught.
Mobile pipeline
The same record-and-compile path for Android over Appium, including the cases where the tappable node and its label are different elements in the view tree.
Trace lookup
Paste a request ID and get the cross-service path it took. Pulled from observability on request, never in the background.
Test case management
Cases tied to the automation that covers them, so a coverage gap is something you can see rather than something you assume.
Runner modes
Five, and they are a frozen contract. Every module has to keep working in all five, because the cost of a runner mode quietly breaking is a suite that looks green while running almost nothing.
- single
- One test, on its own. The loop you use while writing it.
- suite
- One layer or a grouped set, for a targeted regression.
- flow
- Lifecycle and flow tests, where the output of one step feeds the next.
- full
- Every primary layer. The pre-release pass.
- all
- Every module, everything. The nightly.
Rules it will not break
A generated fix is a proposal until a person approves it. Nothing self-applies.
Deterministic checks run before AI, always. AI is the last resort, not the first.
The recording is immutable. Everything else is derived and can be rebuilt from it.
A sandbox pass proves a fix compiles and runs. It does not prove the bug is gone.
Secrets never enter a recording, a report or a repository.
What it is not
A product page without this section is a brochure.
It is internal tooling for one team, not a commercial product. There is no pricing page and no support line.
It assumes a web or Android application it can drive directly. It is not a unit-test runner.
A sandbox-verified fix proves the change compiles and runs. It does not prove the bug is gone, which is why a human still approves it.
How a test gets made
- 1
Record
A real session on the real app becomes a structured recording.
- 2
Compile
The recording compiles to readable Gherkin and a runnable Playwright spec.
- 3
Run
UI, API, flow and security layers execute together across five runner modes.
- 4
Heal
When a locator rots, the run proposes a fix and verifies it in a sandbox.
- 5
Learn
Only verified fixes are kept. Everything else stays advisory until a human approves it.
Outcome
- ✓Investigating a failed run happens in one tool instead of five.
- ✓Recorded sessions become maintained scripts, so coverage grows without hand-writing every selector.
- ✓It is internal tooling rather than a product. I built it because the off-the-shelf options did not fit how the team works.
Stack
- Python 3.12
- Playwright
- asyncio
- FastAPI
- Next.js
- Docker