An open-source locator healing and intent-driven test generation engine for Windows desktop and web.
Automation Sandbox is an open alternative to the black-box locator recovery in commercial tools, centered on a pure-heuristic structural similarity engine (~23ms for 3,000 controls on developer hardware, 0 cost; see SyntheticTreeBenchmarkTests), supplemented by an explainable component scorer, an opt-in multi-provider LLM fallback with an independent-agreement quorum, and an intent-driven test generation pipeline. Desktop support is built on FlaUI (Microsoft UI Automation); web support on the Microsoft.Playwright .NET SDK.
- Deterministic-first. A pure C# structural scorer decides on its own — zero tokens, zero API cost, sub-50ms on a 3,000-control tree. Most healing never touches a model.
- The LLM is never the decision maker. It is an opt-in fallback, only when the heuristic is not confident, and a pick needs an independent-agreement quorum (≥ 2 providers naming the same candidate). Agreement is permission to consider a pick — not evidence it is correct.
-
Nothing is applied silently. The shipped
HealingMode.Reviewdefault changes no locators; a heal commits only in opt-inAutoHeal, and only after the retried action actually succeeds. - PII/secret redaction is on by default before any candidate data reaches an LLM, with prompt-injection hardening on that data. Every decision — which signal contributed what weight, which providers voted, the outcome — is written to a per-decision audit trail (JSON + HTML). See the LLM Security Model.
Yesterday the checkout test stored this locator; today a UI refactor changed both its ID and label:
| Stored snapshot | Live candidate after refactor |
|---|---|
#btn-submit · “Submit order” |
#checkout-confirm · “Confirm order” |
The original lookup fails. SelfHealingEngine captures the live tree, scores candidates from
control type, parent, sibling position, name, and geometry, then retries the same action only
when the confidence, evidence, and ambiguity gates pass:
- click #btn-submit // locator not found
+ click #checkout-confirm // selected candidate; retry succeededOnly that successful retry updates the locator repository. The same decision is written to an auditable healing report; an illustrative excerpt looks like this:
{
"SchemaVersion": 8,
"LocatorKey": "#btn-submit",
"Outcome": "accepted",
"AgreedProviders": [],
"ProviderAttempts": {},
"EvidenceCoverage": 0.85,
"Score": 0.78,
"RunnerUpScore": 0.32,
"ProposedSnapshot": { "AutomationId": "checkout-confirm", "Name": "Confirm order" }
}If the evidence is weak, candidates are tied, or the retried action fails, the engine records that outcome for review and does not persist the proposed locator.
🌐 Try it Live (Web & Playwright): Run the end-to-end browser sample (
dotnet run --project samples/PlaywrightEndToEndQuickstart) demonstrating real Playwright live DOM capture, safe healing of a refactored button, and false-heal prevention on a deleted element across two app versions with interactive HTML reporting. See the Playwright End-to-End Quickstart.
🚀 Published Package Quickstart: Go from
dotnet add package AutomationSandbox.SelfHealing --prereleaseto a successful persisted heal with the Published Package Quickstart and its maintained runnable sample.
🔌 Already have a test suite? Adding Self-Healing to an Existing Test Suite covers the
Observe→AutoHealrollout and minimal wiring for Playwright, NUnit, xUnit, Reqnroll, and FlaUI.
📚 Documentation Hub & GitHub Pages: For complete guides, detailed architecture, JSON schemas, and API references, visit our Documentation Hub.
📦 Preview Packages: The latest prerelease packages are published on GitHub Releases and nuget.org. All seven
AutomationSandbox.*packages are available from nuget.org and as GitHub Release assets. The manual Release workflow publishes through Trusted Publishing (OIDC, without a stored API key); the separate Pack workflow remains artifact-only. See the NuGet Packaging Guide.
🎤 Project Showcase: For a bilingual (EN/TR) architecture presentation and executive summary, see PROJECT_SHOWCASE.md.
| What it is | What it isn't |
|---|---|
A locator healing engine: when an AutomationId or DOM locator breaks, it re-resolves the element from structural evidence and explains the decision component by component. |
Not a monolithic test suite. There is no IDE, visual recorder, execution grid, or scheduler — Ranorex Studio and Tricentis Tosca solve a much wider problem. |
| An intent-driven test generation pipeline: plan steps from a goal, match them to a live page or window, record locators, and emit Playwright or FlaUI test skeletons. | Not a closed runtime or proprietary object repository. Locators live in readable JSON you own; every healing decision is auditable from the recorded report. |
| A modular .NET library set that plugs into the runners you already use (xUnit, NUnit, Playwright, FlaUI). | Not a test framework replacement. It does not run your tests, assert for you, or own your test lifecycle. |
| A deterministic-first design: a zero-token heuristic scorer decides on its own, with an opt-in LLM fallback that is guarded against hallucinated picks. | Not a blind AI wrapper. No screenshots or full DOM dumps are shipped to a model on every step; the LLM sees a bounded top-N shortlist, only when the heuristic is not confident. |
For how this scope and approach compare with Healenium and the commercial healers (Testim, Mabl, Ranorex, Functionize) — including where one of those is the better fit — see How It Compares.
The seven packages follow real dependency boundaries (cross-platform core vs. net48/FlaUI, a Playwright dependency vs. pure DTOs). Install only the ones your scenario needs — dependencies below are pulled transitively.
| I want to… | Install | Notes |
|---|---|---|
| Heal broken locators against a UI tree I capture myself | AutomationSandbox.SelfHealing |
The heuristic scorer, SelfHealingEngine, locator repository, and JSON/HTML reports. Pulls UiModel + LlmHealing transitively. |
| …and let an LLM fallback run when the heuristic isn't confident |
(nothing extra) — register a provider from AutomationSandbox.LlmHealing
|
Ships transitively with SelfHealing; no network calls until you configure a provider (opt-in, quorum-gated). |
| …and capture a live Windows desktop tree (FlaUI / UI Automation) | + AutomationSandbox.Discovery |
Windows-only, targets net48. |
| Capture a web DOM snapshot / get Playwright locator suggestions | AutomationSandbox.WebDiscovery |
Framework-agnostic DOM model and PlaywrightLocatorEmitter. |
| …and launch a browser to capture a page with no hand-written Playwright test | + AutomationSandbox.PlaywrightLiveExploration |
Brings Microsoft.Playwright. |
| Generate Playwright / FlaUI test skeletons from an intent | AutomationSandbox.IntentAutomation |
Intent planning, DOM/desktop matching, locator recording, C#/TypeScript codegen. Pulls WebDiscovery. |
AutomationSandbox.UiModel is the shared DTO layer (UiElementInfo and its JSON serializer). Every package above depends on it transitively — you never install UiModel directly.
New to the library? The Published Package Quickstart walks the SelfHealing path end to end; Adding Self-Healing to an Existing Test Suite covers wiring it into a suite you already have.
| Tier | Surface | Compatibility |
|---|---|---|
| 1 — Stable public API | The core types of the seven AutomationSandbox.* packages (UiElementInfo, SelfHealingEngine, SimilarityWeights, HealResult, ILlmHealingProvider, WebElementInfo, IIntentPlanner, …) |
The committed contract. Breaking changes only on a minor bump, with migration notes in the release notes. |
| 2 — Extensibility points |
ILlmHealingProvider, IHealingReportSink, IIntentPlanner implementations |
SemVer-gated; additive members ship with default implementations. |
| 3 — Internal / experimental |
ScenarioRunner internals, ablation harness, offline research evaluators |
Not a NuGet contract; iterates freely. |
While on 0.x, a minor bump (0.2 → 0.3) may carry a breaking change to Tier 1 (always called out in the release notes); a patch bump (0.2.0 → 0.2.1) never does. Full policy and the 1.0 exit criteria: API Stability & Versioning.
| Feature / Module | Status | Description |
|---|---|---|
| Heuristic Self-Healing | ✅ Implemented | Pure C# structural similarity scoring (O(N) execution, zero-cost, deterministic). |
| Explainable Scoring | ✅ Implemented |
ScoreComponents breakdown (ControlType, Parent, Sibling, Name, Position). |
| Offscreen Rectangle Handling | ✅ Implemented | Dynamic exclusion of unusable (0,0,0,0) bounding boxes from position weights. |
| LLM Fallback & Guard | ✅ Implemented | Gemini, Claude, OpenAI-compatible cloud providers (including Groq, Kimi, OpenRouter, and Cloudflare Workers AI), and offline Ollama behind HttpLlmHealingProvider with LlmProviderFactory auto-discovery and Hallucination Guard. |
| Independent Model Agreement | ✅ Implemented | A quorum rule (MinimumConsensusVotes, default 2), attempt telemetry, nightly multi-model evaluation, and a gating live Groq + Mistral assertion on one known ground-truth candidate. Agreement permits an LLM pick; it is not evidence that the pick is correct. |
| Joint Locator Reconciliation | ✅ Implemented | Opt-in ResolveBatch / ResolveBatchAsync ownership guard prevents independently accepted locators from claiming the same live element; it is a targeted collision guard, not an absence detector. |
| Offline AI Healing (Ollama) | ✅ Implemented | 100% offline, zero-cost local LLM healing with llama3.2 via OllamaHealingProvider. |
High-Level SelfHealingEngine |
✅ Implemented | Configurable HealingMode (Review shipped default, Observe, AutoHeal, FailClosed). In Review (default), proposed candidates are routed to report telemetry without modifying locators or executing retries. In opt-in AutoHeal, repository auto-upsert and action retry only commit upon retry success. |
| Intent-Aware Healing | ✅ Implemented |
TestIntent metadata guiding LLM providers for refactoring-resilient healing. |
| Healing Reports & CI Artifacts | ✅ Implemented | Schema-v8 JSON + HTML telemetry for every resolution attempt, including accepted, ambiguous, ownership-conflict, no-consensus, provider-error, and retry-failed outcomes. AcceptedEvents preserves an accepted-only compatibility view. |
| Synthetic Benchmarks | ✅ Implemented | Pure logic benchmark tests on 3,000+ control trees; core targets netstandard2.0 / net8.0 and runs cross-platform across Windows and Linux CI. |
| WinForms & WPF Live Tests | ✅ Implemented | Real UIA scenario tests against WinFormsApp and WpfApp on Windows CI. |
| Discovery Options & Telemetry | ✅ Implemented |
DiscoveryOptions (MaxDepth, MaxElements, Timeout, CancellationToken, IgnoredFilters). |
| Locator Repository JSON | ✅ Implemented | Versioned repository DTOs/serializer, stable LocatorKey, healing history contract, and thread-safe file locking. |
| Playwright Web Automation | ✅ Implemented |
WebDiscovery DOM snapshot model, Shadow DOM / iframe traversal, PlaywrightApplicationConnector, and Playwright locator emitter. |
| NuGet Preview Packaging | ✅ Implemented | Seven validated AutomationSandbox.* packages with README/license/repository metadata, symbol packages, manual artifact packaging, and GitHub prerelease assets. |
| Published-Package Consumer Sample | ✅ Implemented | Cross-platform, API-key-free quickstart consumes AutomationSandbox.SelfHealing from nuget.org (no project reference), runs a persisted heuristic heal, and is verified from a clean package directory in CI. |
| Playwright End-to-End Sample | ✅ Implemented | Live browser quickstart (samples/PlaywrightEndToEndQuickstart) exercising DOM capture, safe healing, and false-heal avoidance on a real two-version app with HTML report telemetry. |
| Intent-Driven Automation | ✅ Implemented |
AutomationSandbox.IntentAutomation includes intent contracts, both a deterministic and an opt-in LLM-backed (LlmIntentPlanner, guarded with fallback) planner, DOM matching against captured WebDiscovery snapshots, locator recording, Playwright C#/TypeScript generation, intent flow reports, and an end-to-end pipeline API. See Intent-Driven Automation guide. |
| Desktop Intent Automation | ✅ Implemented |
IntentDesktopAutomationPipeline mirrors the web intent pipeline for Windows desktop apps: matches intent steps against a live UiElementInfo tree (IntentDesktopExplorationBridge), records accepted locators, generates an xUnit + FlaUI test skeleton (FlaUiCSharpTestGenerator) built on this project's own Discovery.ApplicationConnector, and emits the same schema-v4 IntentFlowReportDocument (JSON + HTML) as the web pipeline. |
| Live Page Exploration | ✅ Implemented |
PlaywrightLiveExplorer (AutomationSandbox.PlaywrightLiveExploration) launches a browser, navigates to a URL, and captures a WebElementInfo DOM snapshot directly via the Microsoft.Playwright .NET SDK — no hand-written Playwright test, and (deliberately) no Node.js-based MCP server. See why. |
| Organic Benchmark & Calibration | ✅ Implemented | Controlled multi-signal locator ablation on organic application trees (HandBrake 1.8.2), empirical score distribution overlap findings, and threshold trade-off analysis. See Benchmark & Calibration Guide. |
Platform breakdown: Core heuristic engine: cross-platform. Desktop automation: Windows-only (FlaUI). Web automation: cross-platform (Playwright).
The core logic (UiModel, SelfHealing, LlmHealing, WebDiscovery, IntentAutomation, PlaywrightLiveExploration) targets netstandard2.0, .NET 8, and .NET 10 with zero FlaUI/Windows dependency, allowing the heuristic engine, scoring, intent planning, and cross-platform unit tests to execute on Linux, macOS, and Windows. CI runs on a matrix across both Windows (windows-latest for full suite including FlaUI) and Linux (ubuntu-latest for cross-platform core and web suite).
flowchart TB
subgraph STAGE1 ["1. Input and Live Tree Capture"]
direction TB
App["App Under Test (WinForms / WPF)"]
Walker["Discovery Module (UiTreeWalker)"]
Snapshot["Live UI Tree (UiElementInfo)"]
BrokenLoc["Broken Locator (Stale AutomationId)"]
App --> Walker --> Snapshot
end
subgraph STAGE2 ["2. Heuristic Engine (netstandard2.0 / .NET 8 / .NET 10)"]
direction TB
Resolver["SelfHealingResolver"]
Pruner["Candidate Pruner (MinCandidateScore >= 0.05)"]
Scorer["SimilarityScorer (Explainable Scoring)"]
Breakdown["ScoreComponents (Type, Parent, Sibling, Name, Position)"]
Snapshot --> Resolver
BrokenLoc --> Resolver
Resolver --> Pruner --> Scorer --> Breakdown
end
subgraph STAGE3 ["3. Decision and Shortlist"]
direction TB
CheckScore{"Score >= 0.50?"}
ConfidentRes["Heuristic Match"]
Shortlist["Top-N Shortlist Builder (Max 20 Candidates)"]
Breakdown --> CheckScore
CheckScore -->|Yes| ConfidentRes
CheckScore -->|No| Shortlist
end
subgraph STAGE4 ["4. LLM Fallback Chain (Opt-in)"]
direction TB
Eval["LlmHealingEvaluator"]
Providers["Configured Providers (Gemini, Claude, OpenAI, Ollama, ...)"]
Guard["Hallucination Guard (Filter Votes: CandidateId in Shortlist?)"]
Consensus{"Independent Agreement Quorum (>= 2 Votes?)"}
LLMRes["LLM Sourced Match"]
HeuristicFallback["Degrade to Best Heuristic Match"]
Shortlist -->|Prompt ~500 Tokens| Eval
Eval --> Providers
Providers --> Guard
Guard --> Consensus
Consensus -->|"Yes (Agreed)"| LLMRes
Consensus -->|"No (Split / Tie / < 2 Votes)"| HeuristicFallback
end
ConfidentRes --> Output["Final HealResult"]
LLMRes --> Output
HeuristicFallback --> Output
When an AutomationId breaks due to a UI refactor or missing XAML property, SelfHealingResolver executes the following multi-stage pipeline:
sequenceDiagram
participant Test as ScenarioRunner
participant Resolver as SelfHealingResolver
participant Scorer as SimilarityScorer
participant LLM as LlmHealingEvaluator
Test->>Resolver: ResolveAsync(expected, liveTree, providers)
Resolver->>Scorer: ScoreCandidates(expected, liveTree)
Scorer-->>Resolver: List<CandidateScore>
alt Confident (score, evidence and runner-up margin all pass)
Resolver-->>Test: HealResult (Source: Heuristic)
else Not confident - LLM fallback
Resolver->>Resolver: Build Top-N Shortlist (c0, c1, ...)
Resolver->>LLM: EvaluateAsync(expected, Shortlist) - all providers in parallel
LLM-->>Resolver: One LlmHealingResult per provider
Resolver->>Resolver: Drop votes whose CandidateId is not in the shortlist (Hallucination Guard)
alt 2+ providers named the same candidate
Resolver-->>Test: HealResult (Source: LLM, AgreedProviders)
else Split vote, tie, or too few usable votes
Resolver-->>Test: HealResult (Fallback: Heuristic)
end
end
The hallucination guard runs before the vote is counted, so a provider naming a candidate outside its shortlist forfeits only its own vote. Self-reported confidence is recorded but never compared across providers. Independent agreement is the quorum rule that permits a pick; it is not a correctness guarantee. See Independent Model Agreement.
[!CAUTION] DOM/UI text and
TestIntentare untrusted input. The Top-N prompt still sends target metadata plus candidate names and automation IDs to every configured provider; built-in PII/secret redaction is applied by default (opt-out). Structural boundary tags and security directives mitigate prompt injection from that text, but they are not a disclosure control — do not enable cloud providers on screens whose captured fields cannot be disclosed to them. See the LLM Healing Security Model before enabling cloud providers.
SimilarityScorer calculates a weighted sum of independent structural signals and returns a detailed ScoreComponents breakdown:
| Component | Default Weight | Description & Calculation Logic |
|---|---|---|
ControlTypeScore |
0.20 |
1.0 if expected.ControlType == candidate.ControlType, else 0.0. (Weighted, not a hard zero-filter). |
ParentControlTypeScore |
0.20 |
1.0 if parent container ControlType matches, else 0.0. |
SiblingPositionScore |
0.15 |
Proportional index distance: $1.0 - \frac{ |
NameScore |
0.20 |
Levenshtein distance similarity on Name property: |
PositionScore |
0.25 |
Euclidean center-point distance score within PositionToleranceRadius (300px). |
Note
Missing Signal Handling: Every signal is nullable. When both sides lack a signal (empty Name, empty ParentControlType, zero sibling metadata, or an unusable bounding box), that signal scores null — it is excluded from the weighted average entirely, never treated as a perfect 1.0 match. Two elements sharing only ControlType can still reach TotalScore = 1.0, but with EvidenceCoverage of only 0.20.
EvidenceCoverage & MinimumEvidenceWeight: EvidenceCoverage is the fraction of the total signal weight backed by non-null evidence. A heuristic match is IsConfident only when Score >= MinimumConfidence and EvidenceCoverage >= MinimumEvidenceWeight (0.40 by default) — a ControlType-only match is therefore never confident, regardless of its score.
RunnerUpScore & MinimumCandidateMargin: a heuristic match additionally requires best - runnerUp >= MinimumCandidateMargin (0.05 by default). Two near-identical candidates mean "I don't know" — the resolver falls back to LLM/manual review instead of silently picking the tie-break winner. The margin gate does not apply to LLM picks (they use the independent-agreement quorum).
Unusable Rectangle Handling: If a control has a (0,0,0,0) bounding box (e.g. offscreen, unrendered, or collapsed), PositionScore evaluates to null — the same missing-signal rule, so offscreen controls are neither penalized nor erroneously awarded 1.0 center-point matches.
Important
How a heuristic match is accepted vs. how an LLM pick is accepted:
-
MinimumConfidence(0.50): Threshold for accepting a heuristic match before falling back to LLM. -
LLM picks use independent model agreement, not confidence. At least
MinimumConsensusVotesproviders (2 by default) must independently name the same candidate. This is a quorum rule, not evidence of correctness: all 34 unanimous deleted-element verdicts in the measured runs were false heals. Self-reported confidence is recorded but never compared or thresholded — one model's 0.72 and another's 0.95 are not on the same scale. A single configured provider therefore never has its pick accepted, and disagreement (including a tie) falls back to the top heuristic candidate. See docs/llm-providers.md.
using UiModel;
using SelfHealing;
// Load expected locator snapshot
var expected = UiElementSnapshot.FromJson(File.ReadAllText("Snapshots/txtEmail.json"));
// Capture live tree via FlaUI / Discovery
var liveTree = UiTreeWalker.BuildTree(window);
// Run pure heuristic resolution
var result = SelfHealingResolver.Resolve(expected, liveTree);
if (result.IsConfident)
{
Console.WriteLine($"[Healed] Matched '{result.Matched!.AutomationId}' with score {result.Score:F2}");
Console.WriteLine($" Evidence Coverage: {result.EvidenceCoverage:F2}");
Console.WriteLine($" Name Score: {result.ScoreBreakdown?.NameScore}"); // null when both sides lack a Name - no evidence, not a match
Console.WriteLine($" Position Score: {result.ScoreBreakdown?.PositionScore}");
}Configure traversal bounds, timeouts, cancellation tokens, and control filters:
using Discovery;
using System.Threading;
using var connector = ApplicationConnector.Launch(@"C:\apps\MyApp.exe");
var window = connector.GetMainWindow();
var options = new DiscoveryOptions
{
MaxDepth = 15,
MaxElements = 3000,
Timeout = TimeSpan.FromSeconds(5),
IncludeOffscreen = false,
IgnoredControlTypes = new HashSet<string> { "Custom", "ScrollBar" }
};
using var cts = new CancellationTokenSource(TimeSpan.FromSeconds(5));
// Run robust discovery
DiscoveryResult result = UiTreeWalker.Discover(window, options, cts.Token);
Console.WriteLine($"Captured {result.CapturedCount} controls in {result.Elapsed.TotalMilliseconds:F0}ms");
Console.WriteLine($"Visited: {result.VisitedCount}, Skipped: {result.SkippedCount}, Errors: {result.ErrorCount}");
if (result.HitMaxElements) Console.WriteLine("Max element limit reached!");
if (result.TimedOut) Console.WriteLine("Discovery timed out gracefully with partial tree.");
if (result.WasCancelled) Console.WriteLine("Discovery was cancelled gracefully with partial tree.");GetMainWindow() allows 30 seconds by default for a desktop application to become
UIA-ready. It fails immediately if the process exits, and its timeout exception
classifies the failure as slow startup, UIA attachment failure, or ambiguous top-level
windows. If the native main-window handle is temporarily unavailable, the connector can
use the application's sole same-process UIA top-level window; it never guesses when
multiple windows are present. Pass an explicit TimeSpan to override the default.
Note
Discovery Timeout & Filtering Semantics:
-
Timeout: Operates as a best-effort traversal budget between nodes, returning a usable partial tree (TimedOut = true) rather than throwing an exception. -
CancellationToken: Stops traversal gracefully at the next checkpoint and returns partial result withWasCancelled = true. -
IncludeOffscreen: When set tofalse(default), controls with zero bounding boxes(0,0,0,0)orIsOffscreen = trueare skipped and excluded fromnode.Children(SkippedCountincremented). -
IgnoredControlTypes&IgnoredClassNames: Matching descendant nodes and their subtrees are excluded fromnode.Children. The requested root remains the traversal anchor and is never removed by filters (depth > 0). -
SiblingCountIntegrity:SiblingCountis calculated over actual captured non-filtered siblings in the tree, ensuring mathematical consistency forSiblingPositionScore. -
Fail-Fast Validation:UiTreeWalker.Discovervalidates options before traversal, throwingArgumentNullExceptionorArgumentOutOfRangeExceptionfor invalid parameters (MaxDepth < 0,MaxElements < 1,Timeout <= 0).
Customize scoring behavior for applications with unstable positions or static control names:
var customWeights = new SimilarityWeights
{
NameWeight = 0.40, // Give higher priority to Name match
PositionWeight = 0.10, // Reduce sensitivity to layout shifts
PositionToleranceRadius = 500.0, // Expand distance radius for high-res displays
MinimumConfidence = 0.60, // Raise heuristic confidence bar before LLM fallback
MinimumConsensusVotes = 2, // Providers that must agree before an LLM pick is accepted
MinCandidateScore = 0.10, // Aggressively prune low-scoring candidates
MaxCandidatesForLlm = 10, // Limit shortlist size
};
var result = SelfHealingResolver.Resolve(expected, liveTree, customWeights);SimilarityWeights are validated before scoring so persisted/configured values fail fast
when thresholds are outside 0.0..1.0, weights are negative, or the LLM shortlist size is
less than one.
Use a caller-owned key for persisted locators. AutomationId remains part of the
snapshot, but it should not be the repository identity because it may be empty,
duplicated, or stale. LocatorRepository owns a single .locator.json file and
guards its load-modify-save cycle with an exclusive file lock, so concurrent callers
(e.g. parallel test collections healing against the same file) don't race:
var repository = new LocatorRepository("locators.json");
var snapshot = UiElementSnapshot.CaptureFirst(liveTree, node =>
node.ControlType == "Group" && node.Name == "Company");
repository.Upsert("CustomerForm.Company", snapshot!, applicationName: "CustomerApp", platform: "windows-uia");When a heal actually happens, bridge the HealResult into a LocatorHealingHistoryEntry
and pass it to Upsert so the repository keeps an audit trail of what changed and why:
var healResult = SelfHealingResolver.Resolve(staleExpected, liveTree);
if (healResult.IsConfident)
{
var entry = LocatorHealingHistoryEntryFactory.FromHealResult(healResult, previousSnapshot: staleExpected);
repository.Upsert("CustomerForm.Email", healResult.Matched!, entry);
}SelfHealingEngine can emit append-only JSON and HTML reports whenever it accepts a
healed locator. Set SELF_HEALING_REPORT_PATH to enable this without changing test code:
$env:SELF_HEALING_REPORT_PATH = "TestResults/healing-report.json"
dotnet test TestAutomation/ScenarioRunner/ScenarioRunner.csproj --configuration Debug --no-buildBy default, the HTML report is written next to the JSON file as
healing-report.html. Override it with SELF_HEALING_REPORT_HTML_PATH when needed.
Updates to an existing JSON report are committed with an atomic same-directory file
replacement: a failed or interrupted commit leaves the previously recorded history in
place instead of deleting it first. The HTML file is derived output written afterward.
Each report event includes:
LocatorKey-
Source(heuristicor the LLM provider name) -
ReviewStatus(accepted,accepted-with-llm, ormanual-review) -
Score,ConfidenceThreshold,CandidateCount -
PreviousSnapshotandAcceptedSnapshot - LLM fields such as
LlmConfidence,LlmProviderName,LlmReasoning, andAgreedProviders(which providers supplied the agreeing votes) when applicable
GitHub Actions uploads both healing-report.json and healing-report.html as the
self-healing-report artifact when healing events occur during CI.
WebDiscovery maps a Playwright-captured DOM snapshot into the same UiElementInfo
shape used by the desktop engine, so the existing self-healing scorer can work across
web and desktop trees:
using PlaywrightLiveExploration;
using WebDiscovery;
// PlaywrightLiveExplorer owns the browser + capture round-trip (see Quick Start #8 below).
// If you're inside your own Playwright test instead, capture the DOM with:
// var json = await page.EvaluateAsync<string>($"() => JSON.stringify(({PlaywrightDomCaptureScript.JavaScript})())");
// var dom = JsonSerializer.Deserialize<WebElementInfo>(json, new JsonSerializerOptions { PropertyNameCaseInsensitive = true });
// (page.EvaluateAsync<WebElementInfo>(...) directly does NOT work - Playwright's own
// deserializer can't populate UiModel.BoundingRectangle, a readonly struct with no setters.)
var liveTree = WebElementMapper.ToUiElementTree(dom);
var result = SelfHealingResolver.Resolve(expectedWebSnapshot, liveTree);
if (result.IsConfident)
{
var healedDomElement = dom.Children.First(e => e.TestId == result.Matched!.AutomationId);
var suggestions = PlaywrightLocatorEmitter.Suggest(healedDomElement);
Console.WriteLine(suggestions[0].Expression); // page.GetByTestId("...")
}The capture script walks regular DOM children, open Shadow DOM roots, and same-origin
iframe documents (capturing hierarchical FrameAncestry and emitting chained Page.FrameLocator
locators). Offscreen elements retain their geometry, while hidden elements are mapped with a zero
bounding rectangle. For cross-origin iframes (where browser Same-Origin Policy blocks parent
DOM inspection), evaluate PlaywrightDomCaptureScript.JavaScript directly inside the target
IFrame context via frame.EvaluateAsync — see Web Automation Guide for details.
Locator string values are emitted as valid C# source literals: quotes, backslashes, and
CR/LF/tab characters are escaped before the suggestions flow into generated tests. CSS
attribute-string values keep their separate CSS escaping inside that C# literal.
The TypeScript generator decodes those C# literal escapes before re-emitting a locator,
preserving complete accessible names even when they contain double quotes.
IntentAutomationPipeline ties the M6 flow together: plan intent steps, match them
against a WebDiscovery DOM snapshot, record accepted locators, generate Playwright
C# and TypeScript test skeletons, and expose a JSON/HTML-ready intent flow report.
using IntentAutomation;
using UiModel;
using WebDiscovery;
var request = new IntentPlanningRequest
{
Name = "Create customer",
Goal = "Create a customer record with valid email",
TargetUrl = "https://example.test/customers",
TestData = new Dictionary<string, string>
{
["email"] = "jane.doe@example.com",
},
};
// In a Playwright test, capture this with PlaywrightDomCaptureScript.JavaScript.
WebElementInfo dom = CaptureDomSnapshotSomehow();
var repository = new LocatorRepository("web.locators.json");
var pipeline = new IntentAutomationPipeline(options: new IntentAutomationPipelineOptions
{
Recording = new IntentLocatorRecordingOptions { ApplicationName = "CustomerPortal" },
Generation = new PlaywrightCSharpTestGenerationOptions { Namespace = "CustomerPortal.Generated" },
});
var result = pipeline.Run(request, dom, repository);
File.WriteAllText("GeneratedCustomerTest.cs", result.PlaywrightCSharpTestCode);
File.WriteAllText("generated-customer.spec.ts", result.PlaywrightTypeScriptTestCode);
new IntentFlowReportFileSink("intent-flow-report.json").Write(result.Report);By default the pipeline plans steps with DeterministicIntentPlanner, which matches a
fixed vocabulary of verbs (save/submit/create/...) in the goal text. Pass
LlmIntentPlanner instead to plan from the goal's natural language directly - it reads
its API key the same way ClaudeHealingProvider does (ANTHROPIC_API_KEY /
ANTHROPIC_MODEL) and degrades safely back to DeterministicIntentPlanner if the key
is missing, the request fails, or the model's response isn't a well-formed step list:
var pipeline = new IntentAutomationPipeline(planner: new LlmIntentPlanner());Both Web (IntentExplorationBridge) and Desktop (IntentDesktopExplorationBridge) exploration bridges protect against unrelated and ambiguous matches:
-
Common interaction vocabulary: intent plans and all three generators support
Hover,UploadFile,PressKey, target-awareWait, andCheck/Uncheckfor checkboxes and radio buttons alongside navigate/fill/select/click/assert.Waitpolls for the matched target instead of emitting a fixed sleep; web uploads require the DOM snapshot'sinput[type=file]signal, while desktop uploads drive the native file dialog from the matched trigger button.Selectis reserved for real dropdown/select/combobox elements — a radio button matchesCheckbut neverUncheck. -
Explicit text-field precedence:
TargetDescriptionis the authoritative free-text matching input.TestIntentremains business-context metadata andExpectedOutcomeremains reporting/assertion metadata; neither can redirect candidate ranking when the prose conflicts. The currentLocatorKeycontract still contributes a matching hint. -
Semantic Overlap Gate (
MinimumSemanticScore = 0.01): Action compatibility alone (e.g. any button) cannot match an intent step without textual/semantic overlap — unrelated elements (e.g. "Delete customer" matching "Export Report") are flagged withRequiresReview = true. -
Runner-Up Margin Check (
MinimumCandidateMargin = 0.05): Competing candidates within margin$< 0.05$ are marked ambiguous for human review rather than guessing. - Unreviewed Persistence Guard: Steps requiring review are excluded from automatic locator repository persistence by default, while retaining full candidate telemetry and runner-up diagnostics in the intent flow report.
Generated Assert steps are emitted from a structured contract, never from a bare presence check:
-
AssertionKind+ExpectedValueonIntentStep:Visible,NotVisible,TextEquals,TextContains,ValueEquals,UrlEquals,UrlContains. Planners produce the contract; the three generators emit code only from it, so an intent like "Order total should be $125" becomesawait Expect(total).ToHaveTextAsync("$125")instead of a visibility check that passes regardless of the value. -
AssertGenerationMode(Strictby default): when an outcome cannot be mapped to a known kind,Strictemits a review marker (Assert.Inconclusive/test.skip/Assert.True(false, ...)depending on the target framework) rather than silently degrading to a check that always passes.Lenientemits a presence check with a// TODOreview comment;Fallbackemits the presence check alone. -
Conservative derivation:
DeterministicIntentPlanner.DeriveAssertiononly produces a value assertion when the outcome carries a value-shaped token (quoted text, currency, number). Generic phrasing such as "the result is visible" staysVisible— a wrong assertion is worse than a weak one. Defaults ship as estimates and are revisited under benchmark issue #15.
IntentDesktopAutomationPipeline is the Windows desktop counterpart to
IntentAutomationPipeline: it plans intent steps with the same IIntentPlanner, matches
them against a live UiElementInfo tree (as captured by Discovery.UiTreeWalker) instead
of a WebDiscovery DOM snapshot, records accepted locators, and generates an xUnit +
FlaUI test skeleton built on this project's own Discovery.ApplicationConnector.
using IntentAutomation;
using UiModel;
var request = new IntentPlanningRequest
{
Name = "Create customer",
Goal = "Create a customer record with valid email",
TestData = new Dictionary<string, string>
{
["email"] = "jane.doe@example.com",
},
};
UiElementInfo window = UiTreeWalker.BuildTree(connector.GetMainWindow());
var repository = new LocatorRepository("desktop.locators.json");
var pipeline = new IntentDesktopAutomationPipeline(options: new IntentDesktopAutomationPipelineOptions
{
Recording = new IntentDesktopLocatorRecordingOptions { ApplicationName = "CustomerApp" },
Generation = new FlaUiCSharpTestGenerationOptions
{
Namespace = "CustomerApp.Generated",
ApplicationExecutablePath = @"CustomerApp\bin\Debug\net48\CustomerApp.exe",
},
});
var result = pipeline.Run(request, window, repository);
File.WriteAllText("GeneratedCustomerDesktopTest.cs", result.FlaUiCSharpTestCode);Note on Report Parity:
IntentDesktopAutomationPipelineResultproduces bothFlaUiCSharpTestCodeand, in.Report, the same schema-v4IntentFlowReportDocument(JSON + HTML) as the web pipeline —Platform = "desktop", every per-step field populated (#372).
Matching favors AutomationId when the recorded snapshot has one, falling back to Name
and then bare ControlType - the same tiering MainFormScenarioTests uses by hand for
panel1, whose AutomationId is deliberately meaningless. The generated code uses direct
FlaUI locators rather than SelfHealingEngine, matching how PlaywrightCSharpTestGenerator
generates direct Playwright locators for the web pipeline: self-healing is a separate,
already-implemented concern (see Quick Start #5), not
something codegen output should wrap every call in.
PlaywrightLiveExplorer closes the gap the "MCP Exploration" docs originally described as
Planned: it launches a browser, navigates to a URL, and captures a WebElementInfo
snapshot directly via the Microsoft.Playwright .NET SDK — no hand-written Playwright test
required to feed a snapshot into IntentAutomationPipeline, IntentExplorationBridge, or
any of the other Quick Start examples above:
using PlaywrightLiveExploration;
await using var explorer = await PlaywrightLiveExplorer.LaunchAsync();
WebElementInfo dom = await explorer.CaptureAsync("https://example.test/customers");
var pipeline = new IntentAutomationPipeline();
var result = pipeline.Run(request, dom, repository);This deliberately uses the Playwright .NET SDK rather than a real Model Context Protocol
bridge: the canonical Playwright MCP server is a Node.js process, and connecting to it
would have made this the first JavaScript/Node.js runtime dependency in an otherwise pure
C#/.NET codebase (see AGENTS.md). Microsoft.Playwright reaches the same functional
outcome as a fully managed .NET client, no Node.js required at runtime. See
Live Page Exploration for the
full rationale.
using LlmHealing;
using System.Net.Http;
// Option A: Auto-discover all configured providers from environment variables (recommended)
var providers = LlmProviderFactory.CreateConfiguredProviders();
// Option B: Explicit provider instantiation
using var httpClient = new HttpClient();
var manualProviders = new ILlmHealingProvider[]
{
new ClaudeHealingProvider(httpClient),
new GeminiHealingProvider(httpClient),
new OpenAiHealingProvider(httpClient),
new OllamaHealingProvider(httpClient)
};
// Falls back to LLM only if heuristic score < MinimumConfidence (0.50), and accepts the
// LLM's answer only if at least two providers independently pick the same candidate.
// Optional platform ("windows-uia", "web-playwright", etc.) tailors the prompt to the target environment.
var result = await SelfHealingResolver.ResolveAsync(expected, liveTree, providers, platform: "web-playwright");
if (result.Source == HealSource.Llm)
{
Console.WriteLine($"[LLM Healed] {string.Join(" + ", result.AgreedProviders)} agreed on '{result.Matched!.AutomationId}'");
Console.WriteLine($" Reasoning: {result.LlmReasoning}");
}All cloud providers share the HttpLlmHealingProvider base architecture with automatic exponential backoff, per-attempt timeout (15s), and overall operation timeout (35s).
LlmProviderFactory auto-discovers configured models from environment variables:
-
ANTHROPIC_API_KEY(+ANTHROPIC_MODEL)$\rightarrow$ Claude (claude-haiku-4-5-20251001) -
GEMINI_API_KEY(+GEMINI_MODEL)$\rightarrow$ Gemini (gemini-3.6-flash) -
OPENAI_API_KEY(+OPENAI_MODEL,OPENAI_ENDPOINT)$\rightarrow$ OpenAI (gpt-4o-mini) -
GROK_API_KEY+GROK_MODEL(+GROK_ENDPOINT)$\rightarrow$ Grok (both required; no guessed model) -
KIMI_API_KEY+KIMI_MODEL(+KIMI_ENDPOINT)$\rightarrow$ Kimi (both required; no guessed model) -
GROQ_API_KEY+GROQ_MODEL(+GROQ_ENDPOINT)$\rightarrow$ Groq (both required; no guessed model) -
OPENROUTER_API_KEY+OPENROUTER_MODEL(+OPENROUTER_ENDPOINT)$\rightarrow$ OpenRouter (both required; no guessed model) -
CLOUDFLARE_API_TOKEN+CLOUDFLARE_ACCOUNT_ID+CLOUDFLARE_MODEL$\rightarrow$ Cloudflare Workers AI (no guessed model) -
MISTRAL_API_KEY+MISTRAL_MODEL$\rightarrow$ Mistral (both required; no guessed model) -
NVIDIA_API_KEY+NVIDIA_MODEL$\rightarrow$ NVIDIA NIM (both required; no guessed model) -
OLLAMA_CLOUD_API_KEY+OLLAMA_CLOUD_MODEL$\rightarrow$ Ollama Cloud (both required; separate from the local daemon below) -
OLLAMA_HOST/OLLAMA_MODEL/OLLAMA_ENABLED=true$\rightarrow$ Ollama (llama3.2) -
LLM_CUSTOM_PROVIDERSJSON array$\rightarrow$ Custom OpenAI-compatible endpoints (DeepSeek, Cerebras, etc.); every entry requires an explicitName,Endpoint,Model, and API key source. Malformed JSON or a missing endpoint/model is skipped with a credential-safe diagnostic instead of falling back to OpenAI defaults or disabling the built-in providers. Use the three-argumentCreateConfiguredProvidersoverload to route diagnostics to an application logger; the existing overload writes them to standard error.
See docs/llm-providers.md for full configuration and agreement-quorum details, and the LLM Healing Security Model for disclosed fields, provider retention, and report-handling requirements.
When several stale locators are resolved against the same captured tree, use the batch API to prevent two independently accepted heals from taking ownership of one live element:
var batch = await SelfHealingResolver.ResolveBatchAsync(
new[]
{
new BatchHealingRequest("checkout.submit", staleSubmit),
new BatchHealingRequest("checkout.cancel", staleCancel),
},
liveTree,
providers);
foreach (var item in batch.Items.Where(item => item.Result.IsConfident))
{
Console.WriteLine($"{item.Request.LocatorKey} -> {item.CandidateIdentity}");
}The API reconciles only candidates the existing heuristic or LLM agreement-quorum gates already
accepted; it never promotes runner-ups. Candidate ownership uses a snapshot-local tree path,
not AutomationId, so empty and duplicate IDs remain distinguishable. Uncontested false
heals are preserved by design: this is a collision guard, not an absence detector. See the
Joint Locator Reconciliation guide.
The test suite in ScenarioRunner covers all core layers with automated assertions and cross-platform verification:
| Target Component | Covered Behaviors | Test File |
|---|---|---|
| Heuristic Scorer | Structural similarity, weight tuning, unusable (0,0,0,0) bounds |
SelfHealingResolverTests, SelfHealingResolverExplainabilityTests |
| Candidate Pruner | Candidate score filtering (MinCandidateScore), Top-N shortlist assembly |
SelfHealingResolverExplainabilityTests |
| Discovery Robustness |
DiscoveryOptions, DiscoveryResult telemetry, filters and limits, plus actionable application-startup failure classification |
DiscoveryRobustnessTests |
| Locator Repository & Snapshots | Versioned JSON persistence, file locking, LocatorKey stability, UiElementSnapshot round-tripping |
LocatorRepositoryTests, UiElementSnapshotTests |
| Self-Healing Engine & Intent Metadata | Repository auto-upsert, action retry, TestIntent-guided healing, JSON/HTML report emission |
SelfHealingEngineTests, TestIntentHealingTests |
| LLM Providers & Guard | Mocked Anthropic/Gemini/OpenAI/Ollama HTTP responses, Hallucination Guard, and provider resilience: retry on transient 429/5xx, fail-fast on 4xx, Retry-After quota ceiling, per-attempt and total timeout budgets, attempt telemetry |
LlmHealingProviderTests, LlmHealingEvaluationTests, OpenAiAndOllamaHealingProviderTests |
| Independent Model Agreement | Quorum acceptance (not a correctness guarantee), split votes and ties treated as disagreement, single-provider rejection, hallucinated votes dropped before counting, AgreedProviders ordering, LlmProviderFactory discovery, non-confident evaluation fixtures |
SelfHealingResolverTests, ConsensusEvaluationTests, EvaluationScenarios |
| Joint Locator Reconciliation | Stronger ownership, ambiguous ties, empty-ID identity, uncontested limitation, provider partial failure, report telemetry, single-locator compatibility | BatchHealingResolverTests, JointAssignmentGeneralizationTests |
| Real-Tree Locator Ablation | Versioned single- and multi-locator mutation recipes, per-locator ground truth, threshold sweeps, and frozen cross-application joint top-claim evaluation | LocatorAblationTests, ShareXAblationTests, JointAssignmentGeneralizationTests |
| Web Discovery | DOM snapshot mapping, Shadow DOM / iframe traversal, hidden/offscreen handling | WebDiscoveryTests |
| Intent Automation (Web) | Deterministic + LLM-backed planning (with guarded fallback), DOM candidate matching/exploration, locator recording, Playwright C#/TypeScript generation, and flow reports | IntentAutomationPipelineTests, IntentPlannerTests, LlmIntentPlannerTests, IntentExplorationBridgeTests, IntentLocatorRepositoryRecorderTests, IntentFlowReportTests, PlaywrightCSharpTestGeneratorTests, PlaywrightTypeScriptTestGeneratorTests |
| Intent Automation (Desktop) |
UiElementInfo candidate matching/exploration, locator recording, xUnit + FlaUI test skeleton generation, and pipeline orchestration |
IntentDesktopExplorationBridgeTests, IntentDesktopLocatorRepositoryRecorderTests, FlaUiCSharpTestGeneratorTests, IntentDesktopAutomationPipelineTests |
| Synthetic Benchmarks | 3,000+ control tree performance, |
SyntheticTreeBenchmarkTests |
| Live UIA Scenarios | End-to-end FlaUI testing against WinForms (net48) and WPF (net8/net10) apps |
MainFormScenarioTests, WpfMainWindowScenarioTests, EndToEndDemoScenarioTests |
| Live Page Exploration | Real headless-Chromium browser launch, navigation, and DOM capture via PlaywrightLiveExplorer against a local HTML fixture |
PlaywrightLiveExplorerTests |
| CI Coverage Visibility | Separate Windows net48 and Linux net8.0 step summaries, overall and per-assembly rows, missing-report handling, artifact retention |
CoverageSummaryWorkflowTests |
| Package Security Audit |
NuGetAudit build-break gate on High/Critical advisories, vulnerable/outdated-package step summaries on both matrix legs, missing/malformed-report handling |
SecurityAuditWorkflowTests |
To collect cross-platform code coverage (coverage.cobertura.xml):
dotnet test TestAutomation/ScenarioRunner/ScenarioRunner.csproj --collect:"XPlat Code Coverage"CI renders the collected Cobertura data into each matrix job's GitHub Step Summary. Windows
net48 and Linux net8.0 are labelled and reported separately, with overall and per-assembly
line/branch coverage. The Linux leg collects with coverlet (XPlat Code Coverage); the Windows
net48 leg uses Microsoft's dotnet-coverage tool instead, because coverlet 8.0+ no longer
ships a .NET Framework-loadable collector (#289). The figures are visibility aids only: they
are never combined into one headline percentage, published as a badge, or used as a threshold
that can fail the build.
Every restore is checked by MSBuild's NuGetAudit (NuGetAuditMode=all, NuGetAuditLevel=moderate
in Directory.Build.props), including transitive packages. Locally this only prints a warning, so
it never blocks development. In CI (ContinuousIntegrationBuild == true), High/Critical advisories
(NU1903 for a direct package, NU1904 for a transitive one) are promoted to build errors, so a
vulnerable dependency cannot merge to main.
Separately, both the Windows and Linux ci.yml matrix legs run dotnet list package --vulnerable --include-transitive and dotnet list package --outdated, and render the results to that job's
GitHub Step Summary via write-security-audit-summary.ps1
— giving visibility into findings below the error threshold (e.g. Moderate/Low severity) as well.
dotnet list TestAutomation/ScenarioRunner/ScenarioRunner.csproj package --vulnerable --include-transitive
dotnet list TestAutomation/ScenarioRunner/ScenarioRunner.csproj package --outdatedBoth test applications (WinFormsApp and WpfApp) implement the same customer registration form, each intentionally embedding a realistic framework-specific locator issue:
| Application | Problematic Control | Cause / Framework Behavior | How Self-Healing Solves It |
|---|---|---|---|
WinForms (net48) |
panel1 |
WinForms automatically surfaces Control.Name as UIA AutomationId. Auto-generated names (e.g. panel1) are frequently left unrenamed in legacy codebases. |
SelfHealingResolver ignores AutomationId during scoring and matches the panel using parent context, child count, and screen bounding box. |
WPF (net8.0-windows / net10.0-windows) |
CompanyPanel (GroupBox) |
WPF never infers AutomationId from x:Name. Unless AutomationProperties.AutomationId is set explicitly in XAML, AutomationId comes back empty. |
SelfHealingResolver matches CompanyPanel using ControlType.Group, parent/sibling position, and header label text. |
AutomationSandbox.sln
├── WinFormsApp/ .NET Framework 4.8 WinForms application under test
├── WpfApp/ .NET 8 / .NET 10 WPF application under test
├── TestAutomation/
│ ├── UiModel/ Shared UiElementInfo, CandidateScore, ScoreComponents & UiElementSnapshot (netstandard2.0, net8.0, net10.0)
│ ├── Discovery/ Live UI tree walker via FlaUI.Core & FlaUI.UIA3 with DiscoveryOptions/Result (net48)
│ ├── SelfHealing/ Heuristic/batch resolver, explainable scoring & shortlist logic (netstandard2.0, net8.0, net10.0)
│ ├── LlmHealing/ HttpLlmHealingProvider base, LlmProviderFactory, Claude, Gemini, OpenAI-compatible cloud providers (including Cloudflare) & offline Ollama (netstandard2.0, net8.0, net10.0)
│ ├── WebDiscovery/ Playwright DOM snapshot mapping, iframe/shadow DOM capture & locator suggestions (netstandard2.0, net8.0, net10.0)
│ ├── IntentAutomation/ Cross-platform intent pipeline & Playwright/FlaUI test generators (netstandard2.0, net8.0, net10.0)
│ ├── PlaywrightLiveExploration/ Live browser page capture via Microsoft.Playwright .NET SDK (netstandard2.0, net8.0, net10.0)
│ ├── NUnitFixtureTests/ NUnit test fixture helper & consumer test suite (net48, net8.0)
│ └── ScenarioRunner/ xUnit test suite: live UIA, self-healing, web discovery, intent automation & live browser coverage (net48 + net8.0 on Windows, net8.0 on Linux)
└── samples/
├── CalibrationCli/ CLI calibration tool for evaluating dataset ablation benchmarks (.NET 8)
├── HeuristicHealingQuickstart/ Console quickstart validating the published NuGet package (.NET 8)
└── PlaywrightEndToEndQuickstart/ End-to-end web test sample using Playwright (.NET 8)
The core logic operates purely on netstandard2.0 / .NET 8 / .NET 10 in-memory trees without requiring Windows UIA COM hooks.
SyntheticTreeBenchmarkTests exercises the heuristic engine against a synthetic UI tree containing 3,000+ candidate controls:
[Benchmark] 3000 candidates scored in 23ms - best score=1.00, candidateCount=3031.- Execution Scaling: O(N) tree traversal and candidate scoring (indicative ~23ms on developer hardware; execution time is hardware-dependent while candidate counts and score outputs are deterministic).
-
Memory Footprint: Allocation-optimized
Flattenenumeration and fast Levenshtein matrix. - Cross-Platform: Benchmark unit tests run natively on Linux, macOS, and Windows.
While synthetic benchmarks measure tree scaling, self-healing quality must be measured on real, organically evolved applications with known ground truth. Every figure below names the application, the sample size, and the threshold it was measured at — a number without that context is not trustworthy, and this project has retracted headline claims before for exactly that reason.
We benchmark against two real application trees using controlled multi-signal locator ablation: a WPF app (HandBrake 1.8.2, 42 authored locators, 176 scenarios) and a WinForms app (ShareX v21.0.0, 29 authored locators, 131 scenarios), across 5 perturbation tiers — pure rename, text drift, position shift, compound drift, and element removal (see full methodology).
A single application cannot tell "this is how the engine behaves" from "this is how one app's structure happens to behave." Measured on two:
| Metric (default weights, MinimumConfidence = 0.50) | HandBrake | ShareX¹ |
|---|---|---|
| Precision | 84.4% | 73.2% |
| Auto-heal recall | 76.9% | 71.4% |
| False heal rate on removed elements | 40.5% | 57.1% |
| Manual review rate | 30.7% | 26.8% |
¹ ShareX figures exclude 15 DataItem grid-row locators from a settings table that are structurally near-identical to their siblings and correctly decline regardless of threshold — see §8 for why, and for the unfiltered numbers.
The false-heal rate did not improve on a second application — it got worse. No static score threshold separates every relocated control from every deleted one whose neighbour looks structurally similar (HandBrake: false heals on removed elements score 0.665–0.955, true compound drifts score 0.749–0.874 — the distributions overlap). Raising MinimumConfidence trades this down at the cost of recall; see the full threshold sweep for both applications, since the same threshold buys a different result on each (7.6%–9.6% false heals on HandBrake vs. 20%–23% on ShareX at 0.75–0.80).
Four live runs across up to seven independent LLM providers (2026-08-16 to 2026-08-18, n = 133 usable scenarios — full results): agreement separates surviving elements from deleted ones better than any heuristic signal tested (94.5% vs. 43.6% unanimous agreement) — but every unanimous verdict on a deleted element across all four runs (34 of 34) was a false heal, including cases where three independently-sourced model families agreed on the same wrong answer. The useful signal in those rejection cases is provider disagreement, not any model recognising that the element is gone. The shipped agreement quorum therefore limits single-model decisions but does not establish correctness or protect against this false-heal mode; widening the provider pool from 3 to 7 did not reduce the failure rate.
For complete methodologies, component breakdowns, and configuration guidance, see the Benchmark & Calibration Guide. The consensus finding is also written up as a standalone story — Can You Trust an LLM to Fix a Broken Locator?.
graph LR
subgraph PhaseA [Phase A: Core Hardening]
M1[M1: Core Hardening MVP - Implemented]
M2[M2: Discovery Robustness - Implemented]
M3[M3: Persistent Locator Repository - Implemented]
M1 --> M2 --> M3
end
subgraph PhaseB [Phase B: Web Automation & Reporting]
M4[M4: Web Adapter, Reports & Docs - Implemented]
end
subgraph PhaseC [Phase C: Productization]
M5[M5: NuGet Preview Packaging - Implemented]
end
subgraph PhaseD [Phase D: Intent-Driven Automation]
M6[M6: Intent Planner & DOM-Snapshot Matching - Implemented]
end
subgraph PhaseE [Phase E: Beta Correctness & Hardening]
P1[Phase 1: Beta Blockers - Closed]
P2[Phase 2: Beta Hardening - Closed]
P1 --> P2
end
subgraph PhaseF [Phase F: Measurement & Adoption]
P3[Phase 3: Calibration & Multi-Model Data - Closed]
P4[Phase 4: Adoption & Consumer Validation - In Progress]
P3 --> P4
end
M3 --> M4 --> M5 --> M6 --> P1
P2 --> P3
Work is now tracked through GitHub milestones rather than the original M1–M6 sequence:
- Phase 1 — Beta Blockers (closed): the correctness gates that had to exist before anything shipped — exception-scoped healing retry, the evidence gate, the runner-up ambiguity margin, the intent semantic gate, LLM divergence tracking, and structured assertions.
-
Phase 2 — Beta Hardening (closed): the independent-agreement quorum for LLM picks (named
Consensusin the API), provider resilience (retry, backoff, dual timeouts,Retry-Afterquota guard), attempt telemetry, cross-platform Linux CI, and packaging parity across all seven libraries. Shipped asv0.2.0-beta.2. - Phase 3 — Post-Beta Measurement (closed): the heuristic and LLM agreement paths were measured against two real applications (HandBrake, ShareX) and four independent multi-provider runs. The #141/#143 frozen study led to #144's opt-in production batch guard: joint top-claim ownership preserved all 79 correct survivor heals and eliminated all 4 observed collisions, but its explicit limit is that 15 uncontested removed-element false heals remain unchanged. The nightly Groq + Mistral ground-truth scenario remains a failing gate rather than collection-only telemetry.
-
Phase 4 — Adoption & Consumer Validation (in progress): validate the published packages from a consumer's perspective and tighten the operator-facing surface. Delivered so far: runnable published-package consumer samples (heuristic, Playwright end-to-end, calibration CLI) each gated in CI, drop-in xUnit/NUnit test fixtures, explicit
Observe/Review/AutoHeal/FailClosedhealing modes withReviewas the default,Conservative/Balanced/Aggressivethreshold profiles plus a per-applicationcalibratecommand, secret/PII redaction on by default, prompt-injection hardening, and a weekly DORA & engine-reliability visibility report. Linux desktop discovery remains a separate research track under #17 rather than an assumed release commitment.
- Windows (for FlaUI.UIA3 / WinForms live scenario execution).
- .NET SDK 8.0 / .NET SDK 10.0 and .NET Framework 4.8 Developer Pack.
# Build entire solution
dotnet build AutomationSandbox.sln --configuration Debug
# Run all test suites
dotnet test TestAutomation/ScenarioRunner/ScenarioRunner.csproj --configuration Debug --no-build
# Audit packages for security vulnerabilities and outdated versions
dotnet list package --vulnerable --include-transitive
dotnet list package --outdated-
Package Security Auditing: Enforced via MSBuild
NuGetAudit(NU1903/NU1904hard errors in CI). -
Verified GitHub Actions Standard: All workflow actions must target active official releases and are verified by
WorkflowActionVersionTests. - Contribution Guidelines: See CONTRIBUTING.md for strict lifecycle rules (DoR, DoD, Assignee assignment, and GitHub Projects board sync).
This project is licensed under the MIT License.
