SUZENT / 文档
SUZENT 使用手册

Images, video & speech

Generate and edit images, read pictures, make videos, and have the agent speak, plus the models each one needs.

此页面目前提供英文内容,中文翻译尚未完成。

These tools sit in the Creative group of the tool picker. Each one runs on its own specialist model, so assign that model first in Settings → Model Roles (see Model roles).

You want toModel role it usesAsks for approvalSaved to
Generate an imageImage GenerationYesimages folder
Edit an imageImage EditingYesimages folder
Read or describe an imageVisionNoNothing saved
Create a videoVideo generationYes, when it is submittedvideos folder
Speak text aloudNone (System) or TTS (API)Noaudio folder (API only)

Folders are inside the conversation's project workspace, so every conversation in the project can reuse the files. Image generation and editing are on by default; the agent can switch on the others when a task needs them. Approvals follow your permission mode.

Images

Ask for a picture and the agent generates it, up to four at a time. To change an existing image, attach it and describe the edit. The agent can combine up to 16 reference images and use a PNG mask to limit the edit to one area, if your model supports it. Each input image can be up to 20 MB. Results appear in the chat and can be edited again.

Reading images

If your chat model can see images, pictures you attach go straight to it. If it can't, the chat shows a note that the images were not sent, and the agent looks at them with its image-reading tool instead. That tool uses the Vision model, which falls back to Primary when Primary can see images. With neither, the tool reports that no Vision model is set.

Video

Assign a Video generation model, then ask, for example, "Generate an 8-second landscape video of mist moving through a forest." Video is a paid, slow job: the agent submits it, checks back on its progress, and shows the finished video in the chat. To animate a picture, attach it and ask to use it as the reference image; your video model must support this.

Speech

Choose how the agent speaks in Settings → Voice & audio, under Speech engine:

  • System speech (no API) uses the voices installed on the device playing the message. No model or API key is needed, but the audio can't be saved, and it only plays in the Suzent chat window, not in chat apps.
  • API speech uses your TTS model and saves audio you can replay. Pick the model under Speech model on the same page. Voices, formats, speed, and style instructions depend on the provider.

New speech plays automatically. Turn off Automatically play new tool speech if you prefer to press Play yourself.

The Creative group also holds the tool for sending messages through your chat apps, which also needs approval.

在 GitHub 上编辑此页 ↗

本页内容