“I built TraceKit to bridge the gap between brittle rule-based scripts and speculative AI testers. The goal was to build a system where the AI acts as a flexible planner (adapting to UI drift and navigating user journeys) while strict, deterministic Chromium DOM APIs serve as an uncompromising ground truth for test verification.”
— TraceKit Architecture Brief
TraceKit is an autonomous browser testing agent designed to eliminate the fragility of traditional end-to-end test maintenance. By pairing multimodal LLM reasoning with deterministic browser automation and assertions, TraceKit turns plain-English test objectives into verifiable browser interactions, structured execution traces, and inspectable artifact reports. Its foundational architectural thesis is: LLMs decide what to do; deterministic browser assertions decide whether it actually worked.
Traditional end-to-end testing frameworks (Playwright, Cypress, Selenium) suffer from recurring maintenance overhead. Subtle DOM restructuring, renamed CSS classes, or adjusted layouts frequently break hardcoded locators even when application logic remains correct. Teams spend substantial engineering cycles updating brittle scripts. Conversely, relying purely on an LLM to evaluate visual screenshots or DOM states often produces hallucinations—grading a broken checkout as 'passing' because it 'looks plausible.'
TraceKit separates planning from verification: multimodal LLM inference interprets user goals and inspects the accessibility tree to select actions, while Playwright assertions evaluate the live DOM state directly, emitting structured timing, locators, and failure classifications.
Captures clean DOM accessibility snapshot, interactive locators, and viewport screenshot.
LLM analyzes current state against goal and historical step context; selects next logical action.
ActionDispatcher executes typed Pydantic action in Patchright Chromium.
Evaluates deterministic Playwright expect assertion directly against Chromium DOM.
Records step trace, visual evidence, timings, and diagnostic failure categorization.
Decoupled Next.js Frontend & FastAPI Backend
The web interface is built with Next.js 16 (React 19, Tailwind CSS) deployed on Vercel, providing live test run streaming, interactive screenshot lightboxes, runs history, and run logs. All API requests are proxied via Next.js rewrites to a secure Caddy reverse proxy on EC2, which routes traffic to a Dockerized FastAPI application server running Uvicorn.
Patchright Chromium Driver & Automation Engine
Utilizes Patchright (an undetectable Playwright fork) driving Chromium browser sessions via Chrome DevTools Protocol (CDP), avoiding bot-detection barriers on modern web applications and ensuring authentic user emulation.
Deterministic Verification Engine
Outcomes are never validated by model speculation. Conditions (visibility, text values, disabled states, element counts) are verified directly against Chromium DOM APIs using Playwright expect assertions.
Letting an LLM evaluate whether a test passed introduces subjective grading and hallucinations. By delegating verification entirely to deterministic Chromium DOM queries, test pass/fail results remain 100% objective and reproducible.
Eliminated false-positive passes caused by visual hallucination while preserving the flexibility of AI-driven navigation.
Raw DOM dumps contain thousands of lines of styling classes, script tags, and non-interactive wrappers that blow through LLM context windows and degrade locator accuracy.
Reduced prompt token consumption by over 75% and dramatically improved LLM locator resolution speed by surfacing only semantic ARIA roles, names, and actionable elements.
Browser test executions can run for tens of seconds or minutes. Blocking HTTP requests would lead to gateway timeouts on reverse proxies.
FastAPI spawns runs as async background tasks; clients poll or stream progress via SSE while artifacts persist safely to disk.