Show it. Say it.
Your AI gets it.
Record your screen and think out loud. Your coding agent turns it into bug tickets, idea boards and todos.
npx skills add AGIHunt/blurt
The app is the recorder anyone can use. The command adds the skill to Claude Code, Codex and other agents.
macOS 13 or later · right-click → Open the first time · Windows is still being tested
What used to take two days of screenshots, red boxes and spreadsheet rows is now a 30-minute walkthrough.
Eyes for your agent. Two days back for you.
One shortcut, then talk.
Press it anywhere
Drag the area to record, or click a window. Tabs, bookmarks and everything else stay out of the video.
Use it and talk
Jump between topics, correct yourself, it's fine. When you say "this", hold ⌃ and drag to circle it.
Click Finish, get back to work
Your agent transcribes locally, splits the ramble into items, picks the frame for each and finds the code that's probably responsible.
Triage like short videos
One item at a time, all keyboard. Then export, or just say "fix these".
Not just bug reports.
Polish a vibe-coded product
An agent built it overnight; you walk through every page and rant. You get a clean bug list with code pointers, ready for the agent to fix.
Browse, then get the spec new
Building a landing page or a store? Tour reference sites, circle what you like and say why. Your agent writes the requirements and design docs. This page started that way.
Capture ideas while browsing
"I like how this site does onboarding… and this pricing page…" You get an idea board: each idea, why, where it came from, the next step.
Hand off from anyone
PMs, designers, ops, clients: record with the app, no repo needed, and send the video. The developer's agent processes it with the code at hand.
Research and walkthroughs
Competitor tours, UX research, "how this works": notes and findings with the frames to prove them.
Not just web apps
Terminals, TUIs and desktop apps record the same way. For phones use the built-in recorder; for hardware, film it. Hand the video to your agent.
Into the tools you already use.
Jira, Notion, Obsidian too: your agent writes through the CLI or MCP you already have. Or skip the export and let it start fixing.
By default, nothing is uploaded.
Stays on your computer
- Recordings. Saved to the workspace you choose,
~/Blurtby default. - Speech recognition. SenseVoice runs on your machine by default: about 240 MB, fast on any CPU.
- Frames and boxes. Scripts pull them from the video locally.
- The review page. Served from your own machine, for you only.
Leaves only when you choose
- What the agent reads. The transcript and the few frames it picks go to the coding agent you already use, which means its model provider. The video itself is never sent.
- Cloud speech recognition. Only if you set your own Groq, OpenAI or DashScope key.
- Exports. Before anything is written to Lark or GitHub, you confirm where and how many.
- Videos for teammates. You copy, you paste.
No black boxes.
A tool that records your screen should let you see what it does. Scripts do the deterministic work, the model does the judgement, and both are in the repo.
Asked the most.
Does it burn a lot of tokens?
The video is never fed to the model. Speech is transcribed locally, for free; the agent reads the text and looks at a few frames only for the moments that become items. Cost follows how many things you talk about, not how long you record. Two real sessions (Claude Code, Opus 5.5, a production web app):
| recording | speech | items | transcription (local) | recording → review page | new input / output tokens | at API prices |
|---|---|---|---|---|---|---|
| 7.0 min | 2.8 min | 9 | 4 s | ~4 min | ~142k / ~19k | ~$2 |
| 8.6 min | 5.0 min | 20 | 5 s | ~5 min | ~150k / ~20k | ~$2 |
Plus cache reads of the ongoing conversation (~4–5M tokens at 1/20 of the input price, included above).
I don't build web frontends. Is it for me?
Yes. If you can see it on a screen or film it with a phone, you can blurt about it: terminals, desktop apps, phone screen recordings, hardware on camera.
Which languages work?
Chinese, English, Japanese, Korean and Cantonese with the default SenseVoice. German, French, Spanish and the rest use Whisper locally or a cloud key; the first run picks one for the language you speak. Items come back in that language.
Does it run on Windows?
It records: on Windows the recorder is Tk + ffmpeg, started by your agent. There is no tray app yet, the pen is macOS-only for now, and it hasn't been verified on enough real machines, so there is no Windows download button here yet.
Next time you have something to say,
just blurt.
npx skills add AGIHunt/blurt