A voice-to-action API for computer-use agents. Your user talks, the agent moves before they finish the sentence. 15 ms per decision on a Qwen 3.8 27B trained to decide, not to chat.
Model update ~15 msWord to decision ~150 msChain of thought 0 tokens
No chain of thoughtNo "are you sure?"No waiting for silence15 msChanges its mind when you do3¢ a minuteNo chain of thoughtNo "are you sure?"No waiting for silence15 msChanges its mind when you do3¢ a minute
Average agent. Chad agent.
The average agent
Screenshots the page, then thinks about it for a while
Waits for you to stop talking
Writes paragraphs of reasoning to click "Next"
"Just to confirm, did you mean…"
Per-token pricing you need a spreadsheet for
The chad agent
Reads the buttons. Decides in 15 ms.
Moves mid-sentence.
One decision. Zero reasoning tokens.
Changes its mind when you do.
3¢ a minute.
Copy. Paste. Done.
Send the choices on screen as a schema. Stream the mic. Get back only what changed. Your key goes where it says mcc_live_….
# Text in, decision out.
curl https://api.megachadcua.com/v1/decide \
-H "Authorization: Bearer mcc_live_…" \
-H "content-type: application/json" \
-d '{"schema": {"action": {"question": "What should the agent do?",
"options": ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"]}},
"text": "take me to sign in"}'# {"fields": {"action": {"value": "Click Sign in", "confidence": 0.98}}, "latency_ms": 212}
# pip install websockets · voice in, decisions out while they talkimport asyncio, json, websockets
KEY = "mcc_live_…"
SCHEMA = {"action": {"question": "What should the agent do?",
"options": ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"]}}
async def run(mic): # mic: async iterator of 16 kHz mono PCM16 chunksasync with websockets.connect("wss://api.megachadcua.com/v1/listen",
additional_headers={"Authorization": f"Bearer {KEY}"}) as ws:
await ws.send(json.dumps({"type": "start", "schema": SCHEMA,
"audio": {"encoding": "pcm_s16le", "sample_rate": 16000}}))
async def read():
async for m in ws:
ev = json.loads(m)
if ev["type"] == "fields": print(ev["fields"])
reader = asyncio.create_task(read())
async for chunk in mic: await ws.send(chunk)
await ws.send(json.dumps({"type": "end"})); await reader
// Browser: stream the mic, act when it's sure.const ws = new WebSocket("wss://api.megachadcua.com/v1/listen?key=mcc_live_…");
ws.onopen = () => ws.send(JSON.stringify({
type: "start",
schema: { action: { question: "What should the agent do?",
options: ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"] } },
audio: { encoding: "pcm_s16le", sample_rate: 16000 },
}));
ws.onmessage = (e) => {
const ev = JSON.parse(e.data);
if (ev.type === "fields" && ev.fields.action?.confidence > 0.9) agent.do(ev.fields.action.value);
};
mic.onframe = (pcm16) => ws.send(pcm16.buffer); // Int16Array at 16 kHz from an AudioWorklet
Pricing.
3¢/ min
First 30 minutes free. No card.
Billed per second while a session is open.
No seats. No tiers. No sales call.
Stripe invoices you monthly. Cancel whenever.
Questions.
Is it fast?
Yes.
Can it handle "no wait, the other one"?
Yes.The answer flips the moment the words change.
Does it work with my agent?
Yes.If it can read the screen, send the buttons as options.