← All work

Internal tool · Nov 2025 — present

Aegis-QA — the tooling I built

An internal platform I built to manage the QA work itself: triggering runs, API testing, Playwright scripts, SQL queries, Datadog lookups and test-case management with automation-script recording, all in one place.

Role
Creator & maintainer
When
Nov 2025 — present

The problem

The work was spread across a dozen tools. Postman here, a Playwright repo there, SQL in a client, Datadog in a browser tab, test cases in a spreadsheet. Every investigation meant stitching them back together by hand.

What I did

  • One place to trigger and watch runs across UI, API, flow and security layers.
  • A recorder that turns a real session into a reviewable script instead of hand-written selectors.
  • API tests, SQL queries and Datadog lookups available next to the run that needs them.
  • Test-case management tied to the automation that covers it.
  • Self-healing locators. Proposed and sandbox-verified, never applied without a human saying yes.
The Aegis-QA interface: a suite run with one failing payout test, the failure triage that classified and diagnosed it, and the Playwright spec compiled from a recorded session.
A run, the triage that follows a failure, and a recording compiled to a spec. Rendered from the real components with invented data.

Why it exists

QA is my profession. AI is my curiosity. Aegis-QA is what happened when the two met: I did not wait for a vendor to sell me a tool, I learned enough AI to build my own.

That is the short version of how a QA engineer ends up maintaining a platform. The QA judgement is mine. AI is how one person shipped something that would normally need a team, and learning enough of it to do that is the part I would do again.

Try it

Open the full portal walkthrough ↗

A working demonstration, running in your browser. It talks to nothing: the run is on a timer, the failure is scripted, and every value is invented. The real tool drives live payment systems, and no capture of that belongs on a public page.

Run a suite

aegis · demo
mode

5 passed · 1 failed · mode suite

  • checkout.guest_card4.1s
  • checkout.saved_card2.2s
  • refund.partial_amount2.8s
  • payout.idempotency_key9.4s
  • ledger.reconcile_on_read1.2s
  • webhook.replay_is_ignored0.9s

Record a test

aegis · demo

4 of 4 captured

What the recorder saw

  • click · button "Add to basket"
  • fill · field "Email" · acme@example.test
  • click · button "Pay now"
  • expect · text "Order ord_4821 confirmed"

What it compiled to

page.get_by_role("button", name="Add to basket").click()
page.get_by_label("Email").fill(email)
page.get_by_role("button", name="Pay now").click()
expect(page.get_by_text(f"Order {order_id} confirmed")).to_be_visible()

Triage a failure

aegis · demo

✗ payout.idempotency_key

stops here until a human approves

  1. ClassifiedLocator no longer matches. This is not a product defect.
  2. DiagnosedThe Continue button moved inside a dialog in build 218.
  3. Fix proposedget_by_role("button", name="Continue") scoped to the dialog.
  4. Sandbox verifiedRuns green in isolation. That is not proof the bug is gone.
  5. Waiting for a humanNothing is applied until someone approves it.

What it replaced

Before this, the job was five windows. The comparison worth making is not against a product with a pricing page. It is against the stack a QA team actually runs when nobody has built them one.

  1. 01

    Write a test

    Before

    Hand-written selectors, rewritten every time the UI moved.

    With Aegis

    Record the real session. The spec is generated, and regenerable from the recording.

  2. 02

    Run it

    Before

    A framework repo driven from a terminal, results scrolling past in the console.

    With Aegis

    One place, with the run history and its artefacts kept together.

  3. 03

    Check the API

    Before

    Postman in another window, with its own environment files to keep in sync.

    With Aegis

    Collections run beside the UI suite, against the same run record.

  4. 04

    Confirm the data really changed

    Before

    A SQL client, queried by hand, when someone remembered to.

    With Aegis

    Saved queries against the state the run just produced. This is where the UI-says-success-but-the-ledger-disagrees bug gets caught.

  5. 05

    A locator stops matching

    Before

    Find it, guess a new one, hope it holds until next week.

    With Aegis

    A proposed replacement, verified in a sandbox, waiting for a human to approve it.

  6. 06

    A run fails at 2am

    Before

    Read the log, then the trace, then the diff, then ask the developer.

    With Aegis

    Classified and diagnosed on arrival, with the trace for the failed request pulled in beside it.

  7. 07

    Know what is covered

    Before

    A spreadsheet, out of date the day after it was written.

    With Aegis

    Cases tied to the automation that covers them, so a gap is visible rather than assumed.

  8. 08

    Test the Android app

    Before

    A separate project, on a separate framework, that nobody maintained.

    With Aegis

    The same record-and-compile path, over Appium.

What is in it

Session recorder

Records a real session on the real application and keeps the intent, meaning which element, which role and which label, rather than the raw selector that happened to match. The recording is the source of truth; everything downstream is generated from it and can be regenerated.

Compiler

Turns a recording into readable Gherkin and a runnable Playwright spec. Nav clicks, entity IDs and CRUD verbs are recovered from the session, so a generated test reads like one a person wrote.

Execution engine

Runs UI, API, flow and security layers across the five runner modes, headed or headless, locally or on the deployed box, with the run history and its artefacts kept together.

Self-healing locators

When a locator rots, the run proposes a replacement and verifies it in a sandbox before anyone sees it. A sandbox pass is advisory. It never resolves the failure, and never writes itself into the knowledge base on its own. A human approves, or it stays a suggestion.

Failure intelligence

Every failed run is classified, diagnosed, given a fix plan, and where the fix is safe and on an allowlist, executed and verified. Deterministic checks run first and AI runs last, so the cheap and certain answer is never skipped in favour of a guess.

Knowledge graph

A live model of the application built from what has actually been observed running, not from a diagram. It answers which tests cover a change, what a change puts at risk, and which surfaces nothing has ever touched.

API and database verification

Collections and contract checks run beside the UI suite, and saved SQL queries confirm the state a run just produced. This is where the UI-says-success-but-the-ledger-disagrees class of bug gets caught.

Mobile pipeline

The same record-and-compile path for Android over Appium, including the cases where the tappable node and its label are different elements in the view tree.

Trace lookup

Paste a request ID and get the cross-service path it took. Pulled from observability on request, never in the background.

Test case management

Cases tied to the automation that covers them, so a coverage gap is something you can see rather than something you assume.

Runner modes

Five, and they are a frozen contract. Every module has to keep working in all five, because the cost of a runner mode quietly breaking is a suite that looks green while running almost nothing.

single
One test, on its own. The loop you use while writing it.
suite
One layer or a grouped set, for a targeted regression.
flow
Lifecycle and flow tests, where the output of one step feeds the next.
full
Every primary layer. The pre-release pass.
all
Every module, everything. The nightly.

Rules it will not break

  • A generated fix is a proposal until a person approves it. Nothing self-applies.

  • Deterministic checks run before AI, always. AI is the last resort, not the first.

  • The recording is immutable. Everything else is derived and can be rebuilt from it.

  • A sandbox pass proves a fix compiles and runs. It does not prove the bug is gone.

  • Secrets never enter a recording, a report or a repository.

What it is not

A product page without this section is a brochure.

  • It is internal tooling for one team, not a commercial product. There is no pricing page and no support line.

  • It assumes a web or Android application it can drive directly. It is not a unit-test runner.

  • A sandbox-verified fix proves the change compiles and runs. It does not prove the bug is gone, which is why a human still approves it.

How a test gets made

  1. 1

    Record

    A real session on the real app becomes a structured recording.

  2. 2

    Compile

    The recording compiles to readable Gherkin and a runnable Playwright spec.

  3. 3

    Run

    UI, API, flow and security layers execute together across five runner modes.

  4. 4

    Heal

    When a locator rots, the run proposes a fix and verifies it in a sandbox.

  5. 5

    Learn

    Only verified fixes are kept. Everything else stays advisory until a human approves it.

Outcome

  • ✓Investigating a failed run happens in one tool instead of five.
  • ✓Recorded sessions become maintained scripts, so coverage grows without hand-writing every selector.
  • ✓It is internal tooling rather than a product. I built it because the off-the-shelf options did not fit how the team works.

Stack

  • Python 3.12
  • Playwright
  • asyncio
  • FastAPI
  • Next.js
  • Docker