Skip to content
Asia/Kolkata
ProjectsJuly 7, 2026

FlakeIQ — Flake Tracking and LLM-Powered Failure Classification

image
FlakeIQ is a post-run analysis pipeline for mobilewright E2E tests that captures, classifies, and visualizes test flakiness. It consists of a Playwright reporter that records the last 10 pw:api steps per test, an LLM-based classifier that categorizes failures into root cause types, and a self-contained Chart.js dashboard for trend analysis. The pipeline has three stages:
  • Capture — A custom Playwright reporter (flake-reporter.js) intercepts every pw:api step during test execution. On onTestEnd, it writes a JSONL record with the test name, platform, device ID, duration, error message, last 10 actions, and a derived screen name.
  • Classify — classify.py reads the JSONL file. For each failed test, it sends the error message and last actions to Ollama running llama3.2 locally. The LLM returns one of five categories: REAL_BUG, TIMEOUT_FLAKE, DEVICE_FLAKE, LOCATOR_FLAKE, or UNKNOWN, along with a one-sentence reason. Results are upserted into SQLite.
  • Visualize — dashboard.py is a zero-dependency HTTP server that serves a Chart.js dashboard. It exposes 14 API endpoints covering flake rate trends, classification breakdown, action-type analysis, platform comparison, daily volume, duration distribution, a screen × day heatmap, device health, and session-level drill-down.
  • Real-time Session Tracking — Each test run gets a unique session_id timestamp. The dashboard groups runs by session, showing per-test pass/fail badges with durations.
  • LLM Classification — Ollama (llama3.2) classifies each failure into one of five root cause categories. The classifier handles Ollama being offline gracefully (skips with a warning).
  • Screen Name Inference — Parses the "Screen - action" naming convention from test titles to tag every record with a screen name for heatmap aggregation.
  • Seed Demo Mode — seed.py generates 5,100 synthetic records across 30 days with realistic failure patterns, 4 devices, and 11 screens. dashboard.py --seed auto-loads it.
  • Zero External Python Dependencies — All Python scripts use stdlib only: sqlite3, json, http.server, urllib.
mobilewright test
  │
  ▼  (FlakeReporter captures last 10 pw:api steps)
flake-results.jsonl
  │
  ▼  (classify.py + Ollama)
flake.db  ◄── seed.py (synthetic data)
  │
  ▼  (dashboard.py)
HTTP server :8080 → Chart.js dashboard
| Endpoint | Description | |---|---| | /api/stats | Total runs, failures, flake rate, avg duration | | /api/sessions | All test sessions with pass/fail stats | | /api/latest-session | Latest session + per-test results | | /api/flake-rate | Daily flake rate trend line | | /api/breakdown | Classification doughnut chart | | /api/by-action | Flake rate by last action type | | /api/by-platform | Pass/fail per platform | | /api/top-flakes | Flakiest tests ranked by fail rate | | /api/devices | Device-level pass/fail stats | | /api/heatmap | Screen × day flake rate grid |
  • Python 3.13+ — stdlib-only scripts for classification, dashboard, and seed generation
  • Node.js — Playwright reporter (flake-reporter.js)
  • Ollama / llama3.2 — On-device LLM for failure classification
  • SQLite — Local database via Python sqlite3
  • Chart.js — Dashboard charts served via CDN
  • mobilewright — Mobile E2E test framework (upstream integration)
A key challenge was designing the screen name inference to handle multiple naming conventions in test titles. The initial regex nav(\w+) extracted "igate" from "navigate", which was fixed by parsing the "Screen - action" pattern explicitly and mapping aliases. Another learning was making the LLM classification robust: Ollama can be unavailable, slow, or return malformed responses. The classifier handles timeouts, connection errors, and parse failures gracefully, falling back to UNKNOWN without crashing the pipeline. The session tracking required careful reporter design — the Playwright reporter is instantiated per-worker, so a session_id is generated once in the constructor to group all tests from the same run. FlakeIQ provides immediate insight into test health with zero infrastructure cost. The pipeline runs entirely locally with no cloud dependencies. The dashboard has caught real patterns in test execution — timeouts concentrated in fill actions, ANR dialogs on a memory-constrained Android emulator, and visibility issues on scroll-into-view operations — enabling targeted fixes rather than blanket retries.
More work

Related projects

Cross-Platform Mobile E2E Testing with mobilewright

Cross-Platform Mobile E2E Testing with mobilewright

Configured and maintained a 19-test mobile E2E suite using mobilewright covering alerts, animation, calendar, forms, gestures, lists, media, signature, profile, and login flows — all passing on both iOS and Android. Contributed two upstream bug fixes to the mobilewright framework.

Adya — From SwiftUI Prototype to Cross-Platform Rewrite

A minimal daily task manager, first validated as a native SwiftUI/iOS prototype, then rebuilt from scratch in Expo/React Native (New Architecture) to ship one codebase across iOS and Android.

QuantForge — Multi-Agent Market Intelligence & Algorithmic Trading Platform

A multi-agent market intelligence platform for Indian markets (NSE/BSE) built on Zerodha Kite Connect — 13 independent AI agents surface probabilistic signals, with a Risk Agent holding veto power and every strategy required to clear a Backtest → Out-of-Sample → Paper Trading → Risk Review gate before it can touch capital.