Hosted appAgentic dApp builderSui Testnet ↗succeeded

Visual flow builder for Sui protocols that compiles a drag and drop flow into an MCP server or agent skill, with a wallet-free dry run simulation path. Tested in place on Sui Testnet.

Tested in place byDeepQA TeamonSui Testnetatrill.naisu.one/onSep 27, 2026

Run #1model gemini-balanced (vertex)took 10m

8 of 11 scenarios passed, 3 failed, 3 high functional issues after the audit.

Share on X
Rill in the browser during the run

By the numbers

8 of 11
scenarios passed, 3 failed
205
browser actions
42
screenshots
183
model calls
9.8
minutes
11
scenarios
8
passed
3
failed
0
blocked
3
issues
high3

Walkthrough

Every scenario DeepQA drove in the browser, in plan order, with the 35 screenshots it captured along the way. A passing scenario is evidence too.

  1. S1
    Connect test wallet via header

    3 steps, 4 screenshots

    pass
    S1-1.png
    S1, Connect test wallet via header
    S1-3.png
    S1, Connect test wallet via header
    S1-5.png
    S1, Connect test wallet via header
    S1-7.png
    S1, Connect test wallet via header
    • Navigated to home page; network is indicated as testnet in the header logo and footer text.
    • Opened Connect Wallet modal showing detected DeepQA Test Wallet option.
    • Selected DeepQA Test Wallet and verified the header button now displays 'DeepQA Test Wallet', confirming successful connection.
    • Clicked the connected wallet button and verified dropdown menu shows wallet options including Disconnect.
    • Wallet connection succeeded smoothly via the header Connect Wallet button, immediately reflecting 'DeepQA Test Wallet' in the navigation bar.
  2. S2
    Load pre-configured template into canvas

    3 steps, 3 screenshots

    pass
    S2-2.png
    S2, Load pre-configured template into canvas
    S2-4.png
    S2, Load pre-configured template into canvas
    S2-6.png
    S2, Load pre-configured template into canvas
    • Navigated to /builder; verified network is testnet and builder UI is loaded with default Trigger, Capabilities, and Output nodes.
    • Clicked Template button; modal opened showing available presets including 'Swap → Stake'.
    • Loaded 'Swap → Stake' template onto canvas; verified multiple nodes (Trigger, Capabilities, Cetus, Haedal, Output) and connecting edges populated with 2 PTB moves.
    • Successfully navigated to /builder on Sui testnet.
    • Opened the template modal and selected the 'Swap → Stake' preset.
    • The canvas graph accurately populated with Trigger, Capabilities, Cetus swap, Haedal stake, and Output MCP nodes, with edges correctly connected in sequence.
  3. S3
    Simulate PTB flow execution

    3 steps, 4 screenshots

    fail
    S3-2.png
    S3, Simulate PTB flow execution
    S3-4.png
    S3, Simulate PTB flow execution
    S3-6.png
    S3, Simulate PTB flow execution
    S3-10.png
    S3, Simulate PTB flow execution
    • Navigated to /builder, confirmed Sui testnet is active and wallet is connected.
    • Selected the Swap template and populated the flow canvas with Cetus swap action.
    • Navigated to https://rill.naisu.one/builder and verified the network banner displayed 'Rilltestnet' with DeepQA Test Wallet connected.
    • Opened the Template picker and selected the Swap template, replacing the canvas flow with Cetus swap actions.
    • Clicked the Simulate button in the top action bar to trigger PTB simulation.
    • Observed the 'Dry-run & guardrails' simulation modal open, displaying PTB preview step comments, but devInspect stayed on 'devInspect via Rill backend…' without returning gas estimates or balance changes.
  4. S4
    Compile and export MCP Server code

    3 steps, 4 screenshots

    fail
    S4-2.png
    S4, Compile and export MCP Server code
    S4-4.png
    S4, Compile and export MCP Server code
    S4-6.png
    S4, Compile and export MCP Server code
    S4-8.png
    S4, Compile and export MCP Server code
    • Navigated to builder page showing Sui testnet badge and the workflow builder canvas with default graph.
    • Selected the 'Swap → Stake' template, populating the graph with Cetus and Haedal actions wired into an MCP Output node.
    • Navigated to https://rill.naisu.one/builder on Sui testnet.
    • Clicked 'Template' and selected 'Swap → Stake', which populated the canvas with Cetus Swap, Haedal Liquid Stake, and MCP Output nodes.
    • Clicked 'Compile & export', which opened the 'Review & publish' modal with the target flow summary and Publish action.
    • Clicked 'Publish' to generate the MCP Server code / publish the flow.
    • The publication process hangs indefinitely showing 'Publishing flow…' dialog without displaying the generated code or completion screen.
  5. S5
    Introspect protocol with valid package ID

    4 steps, 2 screenshots

    fail
    S5-2.png
    S5, Introspect protocol with valid package ID
    S5-7.png
    S5, Introspect protocol with valid package ID
    • Opened Discover / Import modal and entered Sui framework package ID 0x2.
    • Navigated to https://rill.naisu.one/builder on Sui testnet.
    • Verified the network displayed in the banner was testnet ('Rilltestnet').
    • Clicked the 'Discover / Import' button in the library sidebar to open the on-chain introspection modal.
    • Entered a package ID and clicked the 'Introspect' button.
    • The introspection request stalled indefinitely on 'Reading ABI…' without timing out or displaying an error message or parsed ABI.
  6. S6
    Introspect protocol input validation

    4 steps, 4 screenshots

    pass
    S6-2.png
    S6, Introspect protocol input validation
    S6-4.png
    S6, Introspect protocol input validation
    S6-7.png
    S6, Introspect protocol input validation
    S6-15.png
    S6, Introspect protocol input validation
    • Navigated to https://rill.naisu.one/builder on Sui testnet.
    • Clicked 'Discover / Import' button to open the protocol discovery dialog.
    • Entered invalid package ID 'invalid_package_id' into the package ID input field.
    • Clicked 'Introspect' and observed the application attempt RPC discovery before clearly displaying a failure message ('Failed to fetch') without crashing or hanging indefinitely.
  7. S7
    Filter actions via sidebar search

    3 steps, 2 screenshots

    pass
    S7-2.png
    S7, Filter actions via sidebar search
    S7-5.png
    S7, Filter actions via sidebar search
    • Navigated to /builder; sidebar shows action library with Cetus Swap tokens, Haedal Stake SUI, DeepBook Place limit order and a search input.
    • Typed 'Swap' into the sidebar search input; Haedal and DeepBook actions were filtered out, leaving only Cetus Swap tokens visible.
    • The sidebar search input correctly filters the list of available action cards in real-time when typing 'Swap', showing only Cetus DEX Swap tokens and hiding non-matching cards (Haedal Stake SUI, DeepBook Place limit order).
  8. S8
    Search actions empty state

    2 steps, 2 screenshots

    pass
    S8-2.png
    S8, Search actions empty state
    S8-5.png
    S8, Search actions empty state
    • Navigated to /builder, network shows testnet, and the library sidebar displays actions Cetus, Haedal, DeepBook.
    • Typed 'NonExistentActionString' into the actions search input, and the sidebar displayed the empty state message 'No matching actions.'.
    • The application showed 'testnet' in the header/logo badge.
    • Navigating to /builder loaded the flow builder with the actions library sidebar.
    • Typing 'NonExistentActionString' in the 'Search actions…' input resulted in 'No matching actions.' being displayed cleanly without layout breakdown or errors.
  9. S9
    Open Capabilities configuration panel

    2 steps, 2 screenshots

    pass
    S9-2.png
    S9, Open Capabilities configuration panel
    S9-4.png
    S9, Open Capabilities configuration panel
    • Navigated to /builder, displaying the visual flow canvas with top toolbar buttons including Capabilities2.
    • Clicked the Capabilities button; the Capabilities configuration modal opened showing wallet-level capability manifest, restriction settings, and add restriction options.
    • The application showed network testnet in the header banner.
    • Navigating to /builder loaded the visual flow builder with the toolbar and canvas.
    • Clicking the Capabilities button in the top toolbar opened the Capabilities configuration dialog modal.
    • The modal displays wallet-level capability manifest controls including Budget, Rate limit, editable inputs, and buttons to add further restrictions.
  10. S10
    Auto-arrange graph layout

    3 steps, 4 screenshots

    pass
    S10-2.png
    S10, Auto-arrange graph layout
    S10-4.png
    S10, Auto-arrange graph layout
    S10-6.png
    S10, Auto-arrange graph layout
    S10-8.png
    S10, Auto-arrange graph layout
    • Navigated to /builder with Sui testnet active and initial graph loaded.
    • Selected 'Swap → Stake' template which populated the canvas with nodes and edges.
    • Clicked 'Auto-arrange' button which reorganized the template nodes and edges cleanly across the canvas.
    • Verified builder canvas loads properly on Sui testnet.
    • Opened template dialog and selected 'Swap → Stake' template.
    • Clicked Auto-arrange and confirmed graph layout remained cleanly organized and visible on canvas.
  11. S11
    Verify top-level navigation links

    4 steps, 4 screenshots

    pass
    S11-1.png
    S11, Verify top-level navigation links
    S11-3.png
    S11, Verify top-level navigation links
    S11-7.png
    S11, Verify top-level navigation links
    S11-10.png
    S11, Verify top-level navigation links
    • Navigated to /protocols, displaying flow builder preview and protocol actions like Pyth, Cetus, and Haedal.
    • Navigated to /docs, displaying Getting started documentation with Compose, Configure, Export, and Runtime sections.
    • Navigated to /pitch, displaying pitch presentation deck with slides and navigation controls.
    • Verified that top-level navigation links (Protocols, Docs, Pitch) successfully load their respective pages without 404s, blanks, or errors.

Issues

Findings that survived the Critic's audit. Security-class issues stay summary-only until the maintainers ship a fix.

highconfirmedfunctionalF1 in S3

PTB simulation is unreliable, routinely failing to return gas estimates before a long stall

I navigated to /builder, selected the Swap template, and clicked Simulate. The modal opened but remained stuck on the text 'devInspect via Rill backend…' without showing gas estimates or balance changes, reproducing the issue exactly as reported. Hand check, 2026-09-27: reproduced the same Cetus swap flow in a real browser. The simulate call stalled on 'devInspect via Rill backend…' for about 18 seconds, then failed with "Rill API didn't respond in time — try again", the same message F3's introspection call surfaces. Not an indefinite hang with zero feedback, but a genuinely unreliable backend call, confirmed.

Expected

The simulation modal opens and displays live PTB devInspect execution results including gas estimates and balance changes.

Actual

The simulation modal opens with 'Live simulation · https://api.rill.naisu.one/api' and stalls on 'devInspect via Rill backend…' well past what a user would wait for, then fails with a generic "Rill API didn't respond in time — try again" message instead of showing gas or balance changes.

3 repro steps
  1. Navigate to https://rill.naisu.one/builder
  2. Click the 'Template' button and select the 'Swap' template
  3. Click the 'Simulate' button in the top action bar
highconfirmedfunctionalF2 in S4

Publishing a flow is unreliable, routinely failing after a long stall on 'Publishing flow…'

The publishing process stalls without error feedback recorded in the run. The live replay was inconclusive: the replay ran out of tool calls before it reached the reported state. Hand check, 2026-09-27: reproduced by hand in a real browser, added a Cetus swap node and clicked Publish. The dialog stalled on 'Publishing flow…' for about 18 seconds, then failed with "Rill API didn't respond in time — try again" and returned to the Review & publish screen with Publish re-enabled. Same backend host and same generic error as F1 and F3, not an indefinite hang with zero feedback, but confirmed unreliable.

Expected

The flow should compile and publish or display the generated MCP Server code / export snippets.

Actual

The modal stalls on 'Publishing flow… Publishing action metadata and registering the bounded Rill tools.' well past what a user would wait for, then fails with a generic "Rill API didn't respond in time — try again" message and reverts to the review screen, with no exported code.

5 repro steps
  1. Navigate to https://rill.naisu.one/builder
  2. Click the 'Template' button in the toolbar
  3. Select the 'Swap → Stake' template and confirm replacement
  4. Click the 'Compile & export' button
  5. Click 'Publish' in the Review & publish modal
highconfirmedfunctionalF3 in S5

On-chain protocol introspection is unreliable, routinely failing on a valid package ID after a long stall

The introspection request stalled on 'Reading ABI…' well past what a user would wait for, and there were no prior console errors to attribute this to the test environment. The live replay was inconclusive: the replay ran out of tool calls before it reached the reported state. Hand check, 2026-09-27: repeated twice against the Sui framework package (0x2) in a real browser. Both attempts stalled on 'Reading ABI…' for about 15 seconds, then the app did surface a message, "Rill API didn't respond in time — try again", so this is not an indefinite hang with zero feedback. The underlying defect is real: introspection of a valid, real testnet package consistently fails against Rill's own backend, it just fails with a delayed, generic retry message rather than hanging forever.

Expected

The application should fetch and parse the on-chain package ABI, or display a clear error message/timeout if the package cannot be retrieved.

Actual

The Introspect button transitions to 'Reading ABI…' (disabled) and stalls for an extended period before failing with a generic "API didn't respond in time" message, for a valid, real package ID.

4 repro steps
  1. Navigate to https://rill.naisu.one/builder
  2. Click 'Discover / Import' in the library sidebar
  3. Enter a package ID into the '0x… (the protocol's published package id)' input
  4. Click 'Introspect'

Environment limitations

These failures came from the test environment, not from the application: a credential the sandbox does not hold, a demo nobody may write to, a resource it cannot reach. They are not counted as issues.

  • S6 could not exercise this: Uncaught connection refused error during protocol introspection. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
  • S7 could not exercise this: Uncaught connection refused error during sidebar search. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
  • S8 could not exercise this: Uncaught connection refused error on empty state search. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
  • S9 could not exercise this: Uncaught connection refused error on opening Capabilities panel. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.

Wallet activity

DeepQA injected a test wallet into the browser and recorded every request the app sent to it. Testnet funds only.

App network: testnet

address
0x7b375c…c34f03 ↗
chain
Sui Testnet
browsers opened
3
connects
24
signing requests
0

The app connected the test wallet 24 times and asked for no signature.

Critic audit

An adversarial second pass over every finding before it reaches the report.

3
findings reviewed
3
live replays
0
withdrawn
  • F1confirmed

    Hand check, 2026-09-27: reproduced in a real browser. The simulate call on a Cetus swap template stalled on 'devInspect via Rill backend…' for about 18 seconds, then failed with 'Rill API didn't respond in time — try again', the same message F3's introspection surfaces. Confirmed as an unreliable backend call, not a silent indefinite hang.

  • F2confirmed

    Hand check, 2026-09-27: reproduced in a real browser. Publishing a Cetus swap flow stalled on 'Publishing flow…' for about 18 seconds, then failed with 'Rill API didn't respond in time — try again' and reverted to the review screen. Same backend host and error as F1 and F3, confirmed as unreliable, not a silent indefinite hang.

  • F3confirmed

    Hand check, 2026-09-27: repeated twice in a real browser against the Sui framework package (0x2). Both stalled on 'Reading ABI…' for about 15 seconds, then surfaced 'Rill API didn't respond in time — try again', so this is a slow, unreliable introspection call with a delayed generic error, not an indefinite hang with zero feedback. The underlying failure against a valid package is still real and confirmed.

  • Scenarios S3, S4, and S5 stalled indefinitely on actions without recording console errors, while scenarios S6 through S9 recorded explicit ERR_CONNECTION_REFUSED errors.
  • Hand correction, 2026-09-27: F3's original wording ('hangs indefinitely without error handling or timeout') was checked by hand and is not accurate, the app does surface a generic timeout message after roughly 15 seconds. The underlying introspection failure on a valid package stays confirmed, retitled to describe unreliability rather than a silent infinite hang.
  • Hand correction, 2026-09-27: F1, F2 and F3 were all reproduced by hand in a real browser against api.rill.naisu.one. None hang indefinitely with zero feedback as originally worded, each fails after roughly 15 to 20 seconds with the same generic 'Rill API didn't respond in time — try again' message. All three stay confirmed, the underlying backend unreliability is real, retitled to describe the actual behavior.

Report

QA report: external/rill.naisu.one at hosted

Core backend integrations for transaction simulation, flow publishing, and protocol introspection are unreliable, routinely failing after a long stall.

Testing covered 11 scenarios spanning wallet connection, canvas template loading, graph layout, navigation, protocol introspection, PTB flow simulation, and code compilation and export.

Three primary features were confirmed unreliable against Rill's own backend (api.rill.naisu.one): the PTB flow dry-run simulation, publishing a flow, and protocol introspection. All three were reproduced by hand in a real browser: each stalls for roughly 15 to 20 seconds and then fails with the same generic "Rill API didn't respond in time — try again" message, rather than hanging forever with no feedback as the original scenario runs recorded. The underlying failures are real, just not silent or infinite.

Network connection limits in the test environment prevented backend communication during protocol input validation, action filtering, search empty states, and capabilities configuration, though basic interface behaviors passed. With simulation, publishing, and introspection all failing against the live backend, the primary workflow tools are currently unreliable to the point of being unusable.

Run summary
MetricCount
Scenarios executed11
Passed8
Failed3
Blocked0
Findings raised3
Issues after the audit3
Withdrawn by the audit0
Critical / high / medium / low0 / 3 / 0 / 0

Target: https://rill.naisu.one/ · Testing level: deep_feature · Stack: unknown

Issues
High severity
F1 · PTB simulation is unreliable, routinely failing to return gas estimates before a long stall

Severity: high · Type: functional · Verdict: confirmed · Scenario: S3

I navigated to /builder, selected the Swap template, and clicked Simulate. The modal opened but stalled on the text 'devInspect via Rill backend…' without showing gas estimates or balance changes, reproducing the issue exactly as reported. Hand check, 2026-09-27: reproduced the same Cetus swap flow in a real browser. The simulate call stalled on 'devInspect via Rill backend…' for about 18 seconds, then failed with "Rill API didn't respond in time — try again", the same message F3's introspection call surfaces. Not an indefinite hang with zero feedback, but a genuinely unreliable backend call, confirmed.

Expected: The simulation modal opens and displays live PTB devInspect execution results including gas estimates and balance changes promptly.

Actual: The simulation modal opens with 'Live simulation · https://api.rill.naisu.one/api' and stalls on 'devInspect via Rill backend…' well past what a user would wait for, then fails with a generic "Rill API didn't respond in time — try again" message instead of showing gas or balance changes.

Steps to reproduce:

  1. Navigate to https://rill.naisu.one/builder
  2. Click the 'Template' button and select the 'Swap' template
  3. Click the 'Simulate' button in the top action bar

Evidence: screenshots/S3-10.png

F2 · Publishing a flow is unreliable, routinely failing after a long stall on 'Publishing flow…'

Severity: high · Type: functional · Verdict: confirmed · Scenario: S4

The publishing process stalls without error feedback recorded in the run. The live replay was inconclusive: the replay ran out of tool calls before it reached the reported state. Hand check, 2026-09-27: reproduced by hand in a real browser, added a Cetus swap node and clicked Publish. The dialog stalled on 'Publishing flow…' for about 18 seconds, then failed with "Rill API didn't respond in time — try again" and returned to the Review & publish screen with Publish re-enabled. Same backend host and same generic error as F1 and F3, not an indefinite hang with zero feedback, but confirmed unreliable.

Expected: The flow should compile and publish promptly, or display the generated MCP Server code / export snippets.

Actual: The modal stalls on 'Publishing flow… Publishing action metadata and registering the bounded Rill tools.' well past what a user would wait for, then fails with a generic "Rill API didn't respond in time — try again" message and reverts to the review screen, with no exported code.

Steps to reproduce:

  1. Navigate to https://rill.naisu.one/builder
  2. Click the 'Template' button in the toolbar
  3. Select the 'Swap → Stake' template and confirm replacement
  4. Click the 'Compile & export' button
  5. Click 'Publish' in the Review & publish modal

Evidence: screenshots/S4-8.png

F3 · On-chain protocol introspection is unreliable, routinely failing on a valid package ID after a long stall

Severity: high · Type: functional · Verdict: confirmed · Scenario: S5

The introspection request stalled on 'Reading ABI…' well past what a user would wait for, and there were no prior console errors to attribute this to the test environment. The live replay was inconclusive: the replay ran out of tool calls before it reached the reported state. Hand check, 2026-09-27: repeated twice against the Sui framework package (0x2) in a real browser. Both attempts stalled on 'Reading ABI…' for about 15 seconds, then the app did surface a message, "Rill API didn't respond in time — try again", so this is not an indefinite hang with zero feedback. The underlying defect is real: introspection of a valid, real testnet package consistently fails against Rill's own backend, it just fails with a delayed, generic retry message rather than hanging forever.

Expected: The application should fetch and parse the on-chain package ABI promptly, or display a clear error message quickly if the package cannot be retrieved.

Actual: The Introspect button transitions to 'Reading ABI…' (disabled) and stalls for an extended period before failing with a generic "API didn't respond in time" message, for a valid, real package ID.

Steps to reproduce:

  1. Navigate to https://rill.naisu.one/builder
  2. Click 'Discover / Import' in the library sidebar
  3. Enter a package ID into the '0x… (the protocol's published package id)' input
  4. Click 'Introspect'

Evidence: screenshots/S5-7.png

Environment limitations

These failures came from the test environment, not from the application: a credential the sandbox does not hold, a demo nobody may write to, a resource it cannot reach. They are not counted as issues. They record what this run could not exercise.

  • S6 could not exercise this: Uncaught connection refused error during protocol introspection. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
  • S7 could not exercise this: Uncaught connection refused error during sidebar search. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
  • S8 could not exercise this: Uncaught connection refused error on empty state search. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
  • S9 could not exercise this: Uncaught connection refused error on opening Capabilities panel. A net::ERR_CONNECTION_REFUSED console error is logged. The audit recorded the test environment as the cause, so it is not counted as an issue.
Scenario results
ScenarioPriorityResultIssues
S1 Connect test wallet via headerhighpassnone
S2 Load pre-configured template into canvashighpassnone
S3 Simulate PTB flow executionhighfailF1
S4 Compile and export MCP Server codehighfailF2
S5 Introspect protocol with valid package IDhighfailF3
S6 Introspect protocol input validationmediumpassnone
S7 Filter actions via sidebar searchmediumpassnone
S8 Search actions empty statemediumpassnone
S9 Open Capabilities configuration panelmediumpassnone
S10 Auto-arrange graph layoutmediumpassnone
S11 Verify top-level navigation linkslowpassnone
The audit

The Critic reviewed 3 findings and ran 3 live replays in the browser, each on a fresh page.

  • Scenarios S3, S4, and S5 stalled indefinitely on actions without recording console errors, while scenarios S6 through S9 recorded explicit ERR_CONNECTION_REFUSED errors.
  • Hand correction, 2026-09-27: F1, F2 and F3 were all reproduced by hand in a real browser against api.rill.naisu.one. None hang indefinitely with zero feedback as originally worded, each fails after roughly 15 to 20 seconds with the same generic "Rill API didn't respond in time — try again" message. All three stay confirmed, the underlying backend unreliability is real, retitled to describe the actual behavior.
What to fix first
  1. Make PTB flow simulation reliable and fast, and shorten the long stall before the "didn't respond in time" message (F1).
  2. Make the publishing workflow reliable so action metadata registers and code export completes without a long stall or failure (F2).
  3. Make protocol ABI introspection reliable and fast for a valid package, and shorten the current long stall before the "didn't respond in time" message (F3).
Coverage and caveats

In scope: Sui Wallet connection on testnet; React Flow canvas loading and node templates; PTB dry-run simulation and gas estimations; MCP Server / TypeScript code export generation; On-chain ABI protocol introspection; Builder sidebar search and filtering.

Not covered: Actual on-chain PTB execution (simulation only, out of scope for testnet dry runs); Mainnet capabilities (wallet strictly locked to sui:testnet).

  • The provided test wallet automatically approves connection requests without requiring a manual window switch
  • The target package IDs for introspection tests are valid and published on Sui testnet
  • Graph visual layout changes (auto-arrange) are observable in the DOM or canvas attributes
By the numbers
MetricValue
Scenarios8 passed, 3 failed, 0 blocked of 11 (34 planned steps)
Browser actions205 (46 clicks, 5 inputs, 30 navigations, 124 snapshots)
Screenshots42 (4 explore, 35 scenario, 3 critic), 35 captioned
Coverage5 pages, 2 forms, 5 flows, 0 console errors
Audit3 findings, 3 re-verified live, 3 confirmed, 0 promoted, 0 withdrawn
Model calls183
Tokens940,529 input, 9,600 output, 16,503 thinking
Time10 min
Wallet0 transactions, 0 signatures, 0 refusals on chain sui:testnet
StageCallsInputOutputThinkingSeconds
explore22113,8381,8621,08373
plan14,0691,9321,83330
test136749,3274,1865,693341
critique2371,3911,3537,282139
report11,9042676127

Run log

stagecallstokenstime
Explore22116.8k1m 13s
Plan17.8k30s
Test136759.2k5m 41s
Critique2380k2m 19s
Report12.8k7s
Total183966.6k9m 51s
○Intake
✓Explore
✓Plan
✓Test
✓Critique
✓Report
  • 04:28:15Zexploreexplore started
  • 04:38:06ZexploreExplored / (22 controls, 0 forms)
  • 04:38:06ZexploreExplored /protocols (10 controls, 0 forms)
  • 04:38:06ZexploreExplored /docs (10 controls, 0 forms)
  • 04:38:06ZexploreExplored /pitch (20 controls, 0 forms)
  • 04:38:06ZexploreExplored /builder (26 controls, 0 forms)
  • 04:38:06ZexploreMapped 5 pages, 2 forms, 5 flows in 22 turns.
  • 04:38:06Zexploreexplore completed in 73s.
  • 04:38:06Zplanplan started
  • 04:38:06ZplanPlanned 11 scenarios (5 high, 5 medium, 1 low).
  • 04:38:06Zplanplan completed in 30s.
  • 04:38:06Ztesttest started
  • 04:38:06ZtestS1 executed (pass)
  • 04:38:06ZtestS2 executed (pass)
  • 04:38:06ZtestS3 executed (fail), 1 finding
  • 04:38:06ZtestS4 executed (fail), 1 finding
  • 04:38:06ZtestS5 executed (fail), 1 finding
  • 04:38:06ZtestS6 executed (pass)
  • 04:38:06ZtestS7 executed (pass)
  • 04:38:06ZtestS8 executed (pass)
  • 04:38:06ZtestS9 executed (pass)
  • 04:38:06ZtestS10 executed (pass)
  • 04:38:06ZtestS11 executed (pass)
  • 04:38:06ZtestExecuted 11 scenarios: 8 passed, 3 failed, 0 blocked, 3 findings.
  • 04:38:06Ztesttest completed in 341s.
  • 04:38:06Zcritiquecritique started
  • 04:38:06ZcritiqueReviewed 3 findings; 4 possible defects spotted in passed scenarios.
  • 04:38:06ZcritiqueRe-verified F1: reproduced.
  • 04:38:06ZcritiqueRe-verified F2: inconclusive.
  • 04:38:06ZcritiqueRe-verified F3: inconclusive.
  • 04:38:06ZcritiqueAudit complete: 3 confirmed, 0 withdrawn, 0 promoted, 3 re-verified live.
  • 04:38:06Zcritique4 failures came from the test environment rather than the application. They are reported as environment limitations, not issues.
  • 04:38:06Zcritiquecritique completed in 139s.
  • 04:38:06Zreportreport started
  • 04:38:06ZreportReported 3 issues (0 critical, 3 high, 0 medium, 0 low) from 3 findings.
  • 04:38:06Zreportreport completed in 7s.

Put an agent team on your next pull request.

Connect a repo, dispatch a Run, and read an audited, evidence-backed report the same day.