Notes on voice agents, agentic ops, and shipping AI
Latest writing · approved only
Пока нет одобренных материалов на ленте. Черновики только в Редакции — нажми «Одобрить», и карточка появится здесь.
Model eval as a Slack message: Descript’s Underlord loop on OpenRouter
Descript: Slack → Claude Tag → harness evals in 1–2h → PR. Model swap as a parameter, not a wiring project — same family as shell/Files and Muse Secure VM.
OpenRouter shell + Files API: give any model a sandbox without owning the blast radius
Hosted Linux sandbox for any model — network_policy default deny, file promote, $0.0001/active-sec — same blast-radius family as Muse Secure VM.
OpenMontage + HyperFrames: when agents should write HTML, not fight Remotion
HeyGen HyperFrames as agent-native HTML/CSS/GSAP video; OpenMontage first-class runtime; OpenRouter vision captioning — pick the medium on purpose.
Muse Secure VM: what “agent never sees the password” means for production phone agents
Meta Muse Secure VM / Sentinel / credential surrogates — transfer of trust and blast radius for phone/ops agents that must call CRMs and calendars.
OpenRouter rankings in production: flash defaults, coding-agent traffic, when to escalate
Usage through Sep 20: DeepSeek V4.1 Flash #1 (~15.8T, +219%), GLM 5.3 Flash #2 (~14.1T), Hy4 preview #3 (~12.5T), GPT-5.6 Luna #4 (~9.72T); Hermes/Claude Code/Kilo/Cline heavy — flash default, escalate on purpose.
OpenAI Agents API + GPT-Live: what changes for production phone agents
Full-duplex voice and structured handoffs against the same bars that killed the Telnyx + Vapi-style stack: latency, transfers, naturalness.
I'm ripping out my AI phone stack — here's what actually failed.
Inbound agents on business phone lines looked solved on paper. Latency, transfers, and naturalness said otherwise. Notes from tearing down Telnyx + Vapi and what I'm evaluating next.
Cursor Projects for agentic ops
Short stub: scoped projects as boundaries for agentic work — less chat forever, more shippable ops.