The #1 real-time speech and transcription models purpose-built for voice agents. One API, no tradeoffs between quality and speed.
Today's AI learns from data curated by humans. We're building AI that learns from the world as it is, and gets better with every interaction.
Streaming text-to-speech with natural, expressive voices in 40+ languages.
Real-time transcription tuned for low-latency voice agents.
No tradeoffs between quality and speed, deployed on your terms.
Controls are 36px at 7.2px radius; hero CTAs are 44px at a squarer 4px. Primary is forest green #004e23 with #fefefe text; secondary is canvas with a 1px #e4e3db border and hovers to #f1f0ec. Focus sets border #a1a1a1 plus a 3px ring; disabled drops to opacity 0.5.
There are no real shadows — the only box-shadow is a transparent ring placeholder. Structure is the 1px #e4e3db hairline plus decorative repeating-linear-gradient rule grids and hatch. Motion: house 0.15s cubic-bezier(.4,0,.2,1), expo-out reveals at cubic-bezier(.22,1,.36,1), 18 keyframes including a transcription feed and a four-bar audio meter, and a real prefers-reduced-motion block.