Create audio with AI Studio

Turn canvas material into an explainer, interview, podcast, or story, then review the recording and transcript before you share it.

Some material asks to be read. Other material needs to be heard: the project update someone plays while commuting, the product story a client reviews away from their desk, or the interview-shaped explanation that exposes weak reasoning better than another page of bullets.

Choose Audio file when listening is part of the deliverable. AI Studio turns your prompt and source material into a playable canvas object with a waveform and, when the generated transcript is available, a transcript view for review.

The short version

Open an editable canvas, choose Audio file, and add the prompt and source material. Pick Style and Length, then select Submit. Let the placeholder finish, open the audio preview, and review both Waveform and Transcript before you share the canvas or download the recording.

Choose audio when the work needs a voice

Audio is not a decorative version of a document. It changes how the audience receives the material, so use it when pacing, voice, or a listen-without-looking handoff is valuable.

What you needChooseWhy
A brief people will edit, quote, or scanDocText keeps headings, links, and precise wording easy to revise.
Facts arranged by rows and columnsTableStructure matters more than delivery.
One visual direction or campaign assetImageThe audience needs to see the idea at a glance.
A narrated briefing, conversation, or storyAudio fileVoice and pacing are part of the result.
A result reviewers should open and interact withWeb pageThe layout and page experience matter.

A weekly leadership update is a good audio candidate when the intended audience will listen between meetings. The same update should stay a Doc when the audience needs to approve exact dates, copy passages into a plan, or leave line-by-line edits.

Check access before you start

ItemDetails
Available onEditable, non-embedded canvases in the web app and desktop app.
Who can create audioWorkspace members with edit access to the canvas.
Who cannot create audioGuests, external collaborators, and people with view-only or comment-only access.
What starts the workA successful workspace AI credit check when you submit.
Where the result goesA generated audio object on the current canvas.

If AI Studio is missing, confirm that you opened the canvas itself and have edit access. A workspace member must create the audio for a guest or external reviewer.

Create the recording

  1. Open the canvas where the audio and its source material belong.
  2. Select the notes, documents, images, tables, or earlier results the recording should use.
  3. In AI Studio, choose Audio file.
  4. Write the prompt. Name the audience, purpose, speaker shape, facts that must stay intact, and any terms that need careful pronunciation.
  5. Add files or HTTPS links when the source is not already on the canvas. Wait for every upload to finish.
  6. Choose Style and Length.
  7. Select Submit. AI Studio places a processing object on the canvas while it writes and renders the audio.
  8. Return to the object when playback is ready, then review the recording and transcript.

The title can appear before the playable file is finished. Treat the object as processing until the playback controls are available.

Pick the style by listening job

All five styles use the same audio generation path. Style directs how the material is shaped; it does not switch to a different product or quality tier.

StyleUse it whenWhat to ask for
AutoYou care about the result more than the format.State the audience and outcome, then let AI Studio choose the strongest treatment.
ExplainerA difficult topic needs a clear, concrete walkthrough.Name the listener's starting knowledge, the core question, and the examples that must be included.
InterviewQuestions and answers will make the expertise or disagreement easier to follow.Name the host's role, the guest's point of view, and the questions that should create pressure rather than polite repetition.
PodcastYou want a two-host episode with a hook, conversational rhythm, and a complete arc.Give the episode's thesis, intended listener, and the one idea the ending should land.
StorytellingA sequence, transformation, or case study should unfold as a narrative.Give the setting, turning point, evidence, and ending.

Use Explainer for clarity and Podcast for conversational energy. Use Interview when one voice should ask and the other should carry the expertise. Use Storytelling when order and emotional movement are part of the meaning.

If speaker count matters, write it in the prompt. One presenter and two hosts are clearer instructions than expecting the style name to settle every casting choice.

Treat length as a planning control

LengthBest fit
AutoLet AI Studio read the prompt and source density before deciding the breadth.
ShortA focused update, introduction, clip, or single through-line.
MediumA complete explanation or conversation with room for evidence and examples.
LongA deeper episode with more development, examples, and transitions.

These choices shape the amount of material and the episode arc. They are not hard playback caps. Spoken runtime changes with the source, transcript, speaker cadence, and the final recording.

When an exact target matters, state it in the prompt as well as choosing a length: Create a single-presenter briefing aimed at about 90 seconds. Review the finished runtime before delivery rather than treating the target as a guarantee.

Give the audio a source worth listening to

AI Studio can use the prompt, selected canvas objects, completed attachments, and prompt URLs together. The cleanest recordings start from a small, deliberate source set.

ContextBest use
Selected canvas objectsKeep the recording tied to the notes, table, image, or draft already under review.
Attached filesUse a brief, PDF, transcript, spreadsheet, or other finished source that is not yet on the canvas.
HTTPS linksAdd public material the service can reach. Private, expired, or blocked pages are not reliable source material.
PromptSet the audience, listening situation, structure, speaker count, language, duration target, and facts that must not change.

Do not hand the model six loosely related files and ask for "a podcast." Choose the source that owns the story. Then tell AI Studio what the listener should understand or decide by the end.

For example:

Turn the selected launch brief and risk table into a two-host Explainer for the sales team. Start with the customer problem, cover the three launch risks with their owners, keep every date and metric unchanged, and end with the decision needed this Friday.

That prompt gives the audio an audience, a spine, and a factual boundary. Those three things remove more cleanup than adding mood adjectives.

Let long audio finish in the background

Audio generation can take longer than a document because AI Studio prepares the material, writes the spoken structure, renders the voices, and assembles the playable file. The processing object is the task's live status on the canvas.

You can leave the canvas after the request is accepted. The task continues on the server, so return to the same canvas to review it. Do not submit the same prompt again because the first object is still processing. That creates another credit-spending request, not a faster version of the first one.

Delete the processing object when you mean to abandon the run. Deletion is the cancel gesture, but it is not a promise that provider work already in progress will be reversed or that credits will return. Deleting a finished recording never returns credits.

Listen with the transcript open

Open the finished audio object in preview. The player gives you normal playback controls, playback speed, and a waveform. When transcript text is available, use the Waveform and Transcript tabs to review the same recording in two ways.

Run this review before the object reaches a client or a large team:

  1. Listen to the opening and ending. Confirm the episode makes the promised point rather than merely summarizing the source.
  2. Check names, product terms, dates, figures, and pronunciations against the source.
  3. Read the Transcript for claims that pass too quickly in audio.
  4. Check the pacing at normal speed. Change playback speed for review, not to hide an overlong script.
  5. Confirm the recording contains no private notes, customer data, or internal assumptions that the audience should not hear.

The transcript is a review aid, not a timed subtitle track. Judge the spoken file itself before sharing.

Understand the credit tradeoff

Audio spends workspace AI credits when the generation runs. The final debit reflects the work performed, including the size of the source and the amount of audio processing. Short, Medium, and Long do not display a fixed credit quote in the composer.

Use the shortest format that does the job, and fix a close result manually or through a targeted object-level follow-up instead of regenerating the entire recording for a small wording change. AI credit history is the source of truth after the run.

When audio is not ready or sounds wrong

The object has a title but no playback controls. The audio is still processing. Leave the object in place and return later.

The recording is much longer or shorter than intended. Choose the closest Length, then add a target runtime and a tighter content outline to the prompt. Length directs the plan; it does not trim the final file to an exact second.

The wrong number of speakers appears. State one presenter, two hosts, or the exact interview roles in the prompt.

A name or fact is wrong. Check the transcript against the source. Correct a close result with a focused follow-up or create a new version from cleaner context. Do not approve a polished voice as evidence.

A file was ignored. Confirm the upload finished before submit. Remove failed attachments and attach the source again.

The object stays processing after you return. Refresh the canvas once and keep the original object. If it still does not progress, contact support with the workspace name, canvas link, approximate submit time, output type, prompt summary, and attachment types.

Playback fails. Refresh the preview and confirm the result finished. If the player still reports an error, contact support with the same task details. Do not send private source files unless support asks for the minimum evidence needed.

Give feedback

Was this article helpful?