AIVOOV
AITurn text into natural speech with AIVOOV, a single interface over more than 1500 voices and 140 languages. Agents generate voiceovers, IVR prompts, and audio notifications as a step in a workflow.
What This Integration Enables
AIVOOV is for teams whose audio problem is coverage, not craft. When the requirement is one narrator who sounds like a specific person, that is a different job. When the requirement is the same ninety seconds in forty languages by Friday, breadth is the whole point, and AIVOOV puts more than 1500 voices across more than 140 languages behind a single connector.
An agent picks a voice per language from the catalog, synthesizes plain narration or a full multi-speaker exchange, and gets back a stored file with a downloadable URL that the next step in the flow can attach, upload, or publish. The catalog call is rate limited to twenty requests per day, so a well built flow reads it once and caches the voice identifiers rather than looking them up inside a loop. What that leaves behind is the part worth automating: a rendering step that never becomes a scheduling problem, and a human-in-the-loop pause on the one decision that actually needs ears in the room.
Without FlowRunner
With FlowRunner
Use Case Scenarios
Launch localization across every market at once
A product launch video is approved in English and needs to ship in every market on the list. The agent reads the script and the locale list, resolves a voice per language from the cached catalog, and renders one short sample per locale with Generate Speech. It posts the samples into Slack, grouped by language and addressed to the regional owner who actually speaks it. Each owner keeps the voice or swaps it. Only then does the agent render the full script, upload the finished files to Google Drive in the launch folder, and mark the locale ready in the launch record. The commissioning team stops being the bottleneck on languages they cannot evaluate.
Multi-speaker training and e-learning audio
A compliance training module is written as a dialogue between a manager and an employee. Rather than rendering two files and stitching them, the agent passes an ordered list of voice and text segments to Generate Multi-Voice Speech and gets back a single audio file with each speaker carrying its own voice and delivery settings. When the script changes, the module re-renders in the same flow that flagged the change, so a one line legal correction does not require rebooking anyone.
Audio editions of written content
A newsletter or release note is published. A WordPress publish event starts a flow that summarizes the piece with OpenAI, renders the summary with Generate Speech, and attaches the audio file to the post so readers can listen instead of read. The same pattern serves accessibility requirements: submitted content is read back to the person who submitted it, in the language they submitted it in.
Human-in-Loop Highlight
The moment worth stopping for is the one nobody on the commissioning team can evaluate. Rendering a forty locale batch commits character credits that do not come back, and it ships a voice into markets where no one on the approving side speaks the language. So the agent renders one short sample per locale first, then stops and asks the regional owners directly in Slack: "Here is the Vietnamese sample for the Q3 launch script, voice A, twenty two seconds. Keep this voice, or pick a different one before I render the full three minute script?" The owners answer in their own thread. The agent renders the batch only for the locales that came back approved and holds the rest. The credits get spent on audio somebody signed off, and the person who has to live with the voice in market is the person who chose it.
Agent Capabilities
3 actionsSpeech Synthesis
2- Generate Speech Turn a block of text into spoken audio with a single AIVOOV voice, stored in FlowRunner file storage and returned as a downloadable URL. Pitch, speaking rate, and volume are optional adjustments that default to the voice's own delivery. This is the workhorse call for narration, IVR prompts, and audio notifications.
- Generate Multi-Voice Speech Render an ordered list of voice and text segments into one audio file, with each segment carrying its own voice and delivery settings. Used for dialogue, interviews, character driven e-learning, and any script where two speakers need to sound like two people rather than two files.
Voice Catalog
1- List Voices Retrieve the AIVOOV voice catalog, optionally narrowed to a single language, returning the voice identifier each synthesis call needs. AIVOOV limits this endpoint to twenty calls per day, so read it once at the start of a flow and cache the identifiers rather than calling it per item.
Frequently Asked Questions
What can FlowRunner do with AIVOOV?
FlowRunner agents can run Generate Speech, Generate Multi-Voice Speech, and List Voices in AIVOOV.
Does connecting AIVOOV to FlowRunner require OAuth?
No. AIVOOV connects to FlowRunner with an API key, no OAuth flow required.
Can AIVOOV trigger a FlowRunner workflow automatically?
AIVOOV doesn't currently expose triggers in FlowRunner. It connects as an action step inside workflows started by another trigger.
Start building with AIVOOV
$100 in credits. No card required. Connect in minutes.