FLAGSHIP ARCHITECTURE·Autonomous AI Testing
2025 · PRODUCTION RECORD

TraceKit

Autonomous AI web testing & test orchestration engine

STACK:TypeScriptReact 19Next.js 16Python 3.11FastAPIPatchright (Chromium CDP)Playwright AssertionsDockerCaddy
01
PROBLEM & MOTIVATION
FAILURE MODES & INVARIANTS

“I built TraceKit to bridge the gap between brittle rule-based scripts and speculative AI testers. The goal was to build a system where the AI acts as a flexible planner (adapting to UI drift and navigating user journeys) while strict, deterministic Chromium DOM APIs serve as an uncompromising ground truth for test verification.”

— TraceKit Architecture Brief

TraceKit is an autonomous browser testing agent designed to eliminate the fragility of traditional end-to-end test maintenance. By pairing multimodal LLM reasoning with deterministic browser automation and assertions, TraceKit turns plain-English test objectives into verifiable browser interactions, structured execution traces, and inspectable artifact reports. Its foundational architectural thesis is: LLMs decide what to do; deterministic browser assertions decide whether it actually worked.

THE ARCHITECTURAL FAILURE POINT:

Traditional end-to-end testing frameworks (Playwright, Cypress, Selenium) suffer from recurring maintenance overhead. Subtle DOM restructuring, renamed CSS classes, or adjusted layouts frequently break hardcoded locators even when application logic remains correct. Teams spend substantial engineering cycles updating brittle scripts. Conversely, relying purely on an LLM to evaluate visual screenshots or DOM states often produces hallucinations—grading a broken checkout as 'passing' because it 'looks plausible.'

02
SYSTEM PIPELINE
DETERMINISTIC FLOW
ENGINEERING SOLUTION:

TraceKit separates planning from verification: multimodal LLM inference interprets user goals and inspects the accessibility tree to select actions, while Playwright assertions evaluate the live DOM state directly, emitting structured timing, locators, and failure classifications.

STEP-BY-STEP EXECUTION TRACE:
01OBSERVE

Captures clean DOM accessibility snapshot, interactive locators, and viewport screenshot.

02REASON

LLM analyzes current state against goal and historical step context; selects next logical action.

03ACT

ActionDispatcher executes typed Pydantic action in Patchright Chromium.

04VERIFY

Evaluates deterministic Playwright expect assertion directly against Chromium DOM.

05REPORT

Records step trace, visual evidence, timings, and diagnostic failure categorization.

03
IMPLEMENTATION
TECHNICAL CHAPTERS
CHAPTER 01

Decoupled Next.js Frontend & FastAPI Backend

The web interface is built with Next.js 16 (React 19, Tailwind CSS) deployed on Vercel, providing live test run streaming, interactive screenshot lightboxes, runs history, and run logs. All API requests are proxied via Next.js rewrites to a secure Caddy reverse proxy on EC2, which routes traffic to a Dockerized FastAPI application server running Uvicorn.

—Next.js App Router with real-time SSE streaming for live execution updates
—Caddy reverse proxy handling SSL termination and rate limiting
—FastAPI async RunManager managing background execution lifecycles and user-isolated volume storage (/app/artifacts/runs)
CHAPTER 02

Patchright Chromium Driver & Automation Engine

Utilizes Patchright (an undetectable Playwright fork) driving Chromium browser sessions via Chrome DevTools Protocol (CDP), avoiding bot-detection barriers on modern web applications and ensuring authentic user emulation.

—Isolated browser contexts per test run with clean cookies and local storage
—9 strictly typed Pydantic browser action models with parameter validation
—Automatic waiting and locator retry loops for dynamic SPAs and hydration delays
CHAPTER 03

Deterministic Verification Engine

Outcomes are never validated by model speculation. Conditions (visibility, text values, disabled states, element counts) are verified directly against Chromium DOM APIs using Playwright expect assertions.

—Strict DOM checks: element visibility, exact text match, attribute verification
—Root-cause failure categorization: action failure, locator drift, assertion timeout
—Full trace serialization into inspectable JSON artifacts alongside full-res screenshots
04
TRADEOFFS
ARCHITECTURAL DECISIONS
DECISION 01:Decoupling LLM Planning from DOM Verification
RATIONALE:

Letting an LLM evaluate whether a test passed introduces subjective grading and hallucinations. By delegating verification entirely to deterministic Chromium DOM queries, test pass/fail results remain 100% objective and reproducible.

OUTCOME & ARCHITECTURAL IMPACT:

Eliminated false-positive passes caused by visual hallucination while preserving the flexibility of AI-driven navigation.

DECISION 02:Accessibility Tree Snapshotting over Raw DOM Ingestion
RATIONALE:

Raw DOM dumps contain thousands of lines of styling classes, script tags, and non-interactive wrappers that blow through LLM context windows and degrade locator accuracy.

OUTCOME & ARCHITECTURAL IMPACT:

Reduced prompt token consumption by over 75% and dramatically improved LLM locator resolution speed by surfacing only semantic ARIA roles, names, and actionable elements.

DECISION 03:Asynchronous RunManager with Persistent Volume Storage
RATIONALE:

Browser test executions can run for tens of seconds or minutes. Blocking HTTP requests would lead to gateway timeouts on reverse proxies.

OUTCOME & ARCHITECTURAL IMPACT:

FastAPI spawns runs as async background tasks; clients poll or stream progress via SSE while artifacts persist safely to disk.

05
RESULT
VALIDATION & LESSONS
KEY LESSONS LEARNED:
✓Deep understanding of the Chrome DevTools Protocol (CDP) and headless browser automation internals.
✓How to architect resilient agentic loops that balance generative AI planning with deterministic software engineering assertions.
✓Managing production container lifecycles, reverse proxies, and async background task scheduling in FastAPI.
FUTURE HORIZON:
—Support for concurrent multi-browser cross-platform test matrix execution (Firefox, WebKit).
—Visual regression diffing engine comparing viewport baseline snapshots.
—Export capability allowing autonomous test traces to be exported directly as idiomatic Playwright test scripts.
HIMANSHU PATRO

Building systems. Learning in public.
Based in Jamshedpur, Jharkhand, India.

© 2026 HIMANSHU PATRO