Knowa

Show it your screen. Talk it through. Get back the finished thing.

Knowa is a screen-and-voice assistant. Point it at whatever is on your screen β€” a spreadsheet, a dashboard, a tangled email thread β€” say what you want out of it, and Knowa hands back a finished artifact instead of a set of instructions.

Try it now

No sign-up, no waitlist β€” an interactive preview runs right in your browser.

How it works

1. Share your screen

Pick a tab, a window, or your whole screen. Whatever you are looking at, Knowa sees.

2. Talk it through

Say what you are looking at and what you want out of it. No forms, no prompts to engineer.

3. Get an artifact

A clean summary, table, or draft you can copy and use β€” the finished thing, not instructions.

What you can make

  • A status update pulled straight from a project board
  • The three numbers that matter, out of a messy spreadsheet
  • Action items lifted from a long email thread
  • A tighter rewrite of a paragraph that is fighting you

Why it’s different

Most assistants make you do the describing. You copy things over, explain the context, write a careful prompt β€” and get back advice: steps to follow, suggestions to apply. Turning that into the thing you actually needed is still your job.

Knowa flips that. It already sees what you see, so you skip the describing entirely. You talk the way you’d talk to a colleague looking over your shoulder, and what comes back is the deliverable itself β€” ready to paste.

Under the hood

The proof of concept is a deliberately thin pipeline: the browser captures screen frames and voice, speech is transcribed with Whisper, and both go to Claude’s vision model, which writes the artifact as markdown. No agent gymnastics β€” the bet is that seeing the screen plus hearing intent is enough context to do the job in one shot.

Questions

Is this a real product?

It is a working proof of concept. The capture pipeline runs end-to-end privately; the preview on this site is simulated β€” same flow, canned output β€” so it works with no permissions, sign-up, or API keys.

What happens to my screen?

In the preview, nothing β€” no screen share ever starts. The POC is built to use frames only for the duration of a capture, to produce your artifact, and to keep nothing afterward.

Why voice instead of typing?

Talking is faster and looser than writing a prompt. You can point at things β€” "this column," "that thread" β€” the way you would with a colleague looking over your shoulder.

What is next?

A live mode where you can ask follow-ups without re-sharing, and artifact types beyond markdown β€” slides, emails, spreadsheets.

Screen + voice Β· proof of concept

Try it now