Explore

Public Runs on real apps.

Only published results appear here. DeepQA runs far more than it publishes — internal benchmark campaigns and private repos stay private, and publishing is an explicit choice. Every Run below was executed and published by the DeepQA team: scenarios driven in a real browser, findings adversarially audited by the Critic, and a report you can read end to end.

124
published runs
1463
scenarios executed
255
issues confirmed
15
criticals found
118
apps published
18.5k
browser actions
4.1k
screenshots
135M
tokens

Campaigns

A campaign puts a group of apps through the whole pipeline and publishes every report as it lands.

Arc Campaign

Arc testnet

Celebrating Circle's Arc mainnet launch: DeepQA tests Arc ecosystem apps in place, with an injected test wallet on Arc testnet.

54
apps
627
scenarios
8.2k
browser actions
9.9k
model calls
66.9M
tokens
12 h
testing time

GitHub Agent Apps Campaign

Open source

DeepQA tests agent apps that are trending on GitHub: workflow builders, chat UIs, coding agents, research agents and observability tools.

56
apps
631
scenarios
9.6k
browser actions
9.1k
model calls
57.2M
tokens
8.5 h
testing time

124 runs across 118 apps (showing 1 to 12)

subtitle-translator in the browser during a run
rockbenben/subtitle-translatorLLM translation toolSandbox, from source

Open-source batch subtitle translator with many LLM and translation providers, multi-language output and provider settings. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
m3e-canvas in the browser during a run
lnkiai/m3e-canvasDesign canvas for agent promptsSandbox, from source

Open-source design canvas that turns Material 3 Expressive screens into prompts for coding agents: screens, layers, themes, preview and prompt export. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
ChatChat in the browser during a run
okisdev/ChatChatChat UISandbox, from source

Open-source unified chat and search platform across model providers with sessions, system prompts and provider settings. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
OpenClaw-bot-review in the browser during a run
xmanrui/OpenClaw-bot-reviewAgent fleet dashboardSandbox, from source

Open-source dashboard for a fleet of OpenClaw agents: bot wall, models, sessions, statistics and an alert rule center. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
claudecodeui in the browser during a run
siteboon/claudecodeuiCoding agent web UISandbox, from source

Open-source web and mobile UI for coding agent CLIs: setup wizard, project workspaces, chat, shell, files and source control views. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
crit.md in the browser during a run
Critcrit.mdOpen app ↗tomasz-tomczyk/critHosted appAgent feedback toolHosted, in place

Open-source review tool for coding agents: comment on lines, reply in threads and turn the review into an agent prompt. The public site carries an interactive review demo. Tested in place without signing in.

runs by DeepQA Team
ggml-org-gguf-my-repo.hf.space in the browser during a run
GGUF My Repoggml-org-gguf-my-repo.hf.spaceOpen app ↗ggml-org/llama.cppHosted appModel conversion toolHosted, in place

Hugging Face Space by the llama.cpp organization that converts and quantizes a model repository to GGUF. Tested in place without signing in, so the conversion itself was not run.

runs by DeepQA Team
qwen-qwen3-demo.hf.space in the browser during a run
Qwen3 demoqwen-qwen3-demo.hf.spaceOpen app ↗QwenLM/Qwen3Hosted appLLM chat playgroundHosted, in place

Official Hugging Face Space of Qwen3 by the Qwen team: chat with a thinking mode toggle, thinking budget and conversation history. Tested in place.

runs by DeepQA Team
agent-flow in the browser during a run
patoles/agent-flowAgent run visualizerSandbox, from source

Open-source real-time visualizer for coding agent sessions: agent graph on a canvas, timeline, transcript, file attention and cost overlay. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
neo.u14.app in the browser during a run
Neo Chatneo.u14.appOpen app ↗u14app/neo-chatHosted appChat workspaceHosted, in place

Open-source local-first AI chat workspace with assistants, skills, plugins, knowledge bases, workspaces and MCP servers. Tested in place on its hosted instance.

runs by DeepQA Team
logocreator in the browser during a run
Nutlope/logocreatorAI logo generatorSandbox, from source

Open-source AI logo generator with styles, colors, brand kits and a bring-your-own-key flow. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team
gpt_image_playground in the browser during a run
CookSleep/gpt_image_playgroundImage generation playgroundSandbox, from source

Open-source bring-your-own-key playground for GPT image models with a gallery, favorite folders, search and API settings. Run from its repository in the DeepQA sandbox.

runs by DeepQA Team

Put an agent team on your next pull request.

Connect a repo, dispatch a Run, and read an audited, evidence-backed report the same day.