Kotonia
ログイン今すぐ始める

CHARACTER VOICE CHAT

Free Voice Chat with AI Characters

Pick a public persona and chat with them in text or voice in real time. Japanese, English, and Chinese TTS engines and lipsync avatar display are supported.

Sample lipsync avatar display
SPEAKING

Hi! How can I help you today?

English / VoxCPM2 · Ditto lipsync

A roleplay environment built on real-time voice + AI

Kotonia's character voice chat runs on a real-time VAD + STT + LLM + TTS pipeline. Speak into the mic and the AI replies in voice instantly, with synced lipsync avatars for personas that have an avatar registered. Optimized for language practice, roleplay, casual conversation, and emotional companionship — building relationships through voice.

11 langMultilingual TTS
Free tier100 exchanges/day
LipsyncAvatar sync display

Real-time multilingual voice chat

High-quality multilingual TTS for Japanese, English, Chinese, and more (VoxCPM2). Speak and the pipeline runs STT → AI → TTS back to you instantly.

Public personas + your own

Talk to personas published by other users. You can also save and grow your own personas.

Lipsync avatar display

Register an avatar image to a persona, and Ditto syncs the character's mouth movements to the generated speech in real time.

Saved conversations

Sessions persist with full history. Reference past exchanges and grow the relationship over time.

Web search & code execution, automatically

For tricky questions or anything needing fresh information, the character runs a web search or executes code behind the scenes before replying. Also supports image/video generation and camera/screen-share vision.

Language partnerRoleplay scenariosCasual conversationEmotional companionVoice-input draftingPresentation rehearsal

Basic usage

1

Pick a persona

Choose from public personas — language tutors, casual chat partners, roleplay characters, and more.

2

Grant mic permission

The browser will ask for microphone access on first use. Once granted, subsequent sessions start with one click.

3

Speak into the mic

Press the mic button and talk. VAD auto-detects end of utterance and the AI replies in voice instantly.

4

Continue across sessions

Conversations save automatically per session. Pick up where you left off and the AI remembers what was discussed.

What each setting does

The panel has four tabs — Modes, Character, Memory, and Advanced — and each tab is further divided into a few setting groups. So you can see "which group, and which feature within it," tap a tab below to open it (the order below matches the actual panel). When in doubt, the defaults are fine.

How to open the settings panel

Tap the gear icon (⚙) in the top-right of the Character Chat screen to open this settings panel. The "?" next to it is Help (the how-to guide).

Kotona
* Illustration of the actual screen (for reference)Tap here
Modes tabThe main tab for switching in-conversation behavior
Conversation
Mute / unmute
Stops or restores playback of the AI voice.Use it when you want text only — late at night, or out in public.
History
Opens the list of past conversation sessions and resumes from where you left off.
Clear conversation
Resets the current conversation and starts over.
LLM tier
Tier (Basic / Standard / Premium)
Selects the tier of AI model that composes replies. Basic is free, local, and fastest; Standard is a cost-efficient cloud model; Premium lets you freely pick any frontier model (Grok / Gemini / GPT / DeepSeek, etc.).When unsure, Basic is enough. Standard and above handle emotional nuance and complex topics better. Premium requires a paid plan, or is used when authoring the model for your own character.
Model picker
Appears when Premium is selected; pick the model to use from the frontier-model list.
Custom model ID
Specify a model not in the list by its ID directly (optional, for advanced users).
Modes
Realtime
When ON, you can cut in and speak while the AI is still talking (barge-in). When OFF, you wait for the AI to finish before speaking.Turn it ON for a natural conversational rhythm.
Vision
When ON, shows your camera or shared screen to the AI so it can talk about what it sees. Input source, frame size, and observe interval appear alongside.Input source switches between camera and screen. Frame size (320×180–1280×720): higher is clearer but uses more bandwidth and load (640×360 is plenty). Observe interval (3–30s) is how long it waits after going idle before it looks at the screen.
Auto-talk
When ON, the AI speaks up on its own even when you stay silent. Adjust the frequency with the interval (5–60s).Behaviors like read-aloud, monologue, or small talk should be specified in the system prompt. Works only after the mic is started.
Avatar
When ON, lip-syncs a single registered image with Ditto so its mouth moves in time with the speech.Register a "Ditto source" via the heart icon on a persona image. The free tier allows 12 avatar turns per day; after that the conversation continues as plain voice.
Agent
When ON, hard questions or lookups are answered after running a web search or executing code behind the scenes. When OFF, it stays a plain conversation.
Narration
When ON, the character weaves *stage-direction* prose into replies. That prose is shown on screen only and is not voiced (TTS).Good for novel-style roleplay.
Haptics
When ON, connects a Lovense-compatible device and lets the character drive it in time with the conversation. Available on the Chat plan and above.Pair by scanning the QR code with the Lovense Connect / Remote app. If you do not have a device, leave it OFF.
Image generation model
Image generation model
Controls whether images are generated during the conversation. Choose "None" or "Local GPU (HiDream)".The LoRA stack adds a visual style. Leave it empty and the LLM picks one dynamically (default kotonia03); select one and that stack is always applied. Weight (0–1.5) tunes the strength.
Video generation model
Video generation model
Controls whether a short video is made from a registered image. Choose "None" or "WAN 2.5 (image→video)", then set resolution (480p / 720p / 1080p) and duration (3–10s).Generated from the "latest image" in the conversation.
Character tabShown only for characters you created yourself
Profile
Name / use case / description / tags
Defines your character's basic info, shown in the listing and ranking when published.
Opening line / start scene / reply suggestions
Sets up the conversation entry: the first line on open, the starting scene, and the initial reply buttons.
Publishing & media
Monetization when public
Chooses how others are billed when they use your published character (free / per-message / subscription-only).
Media
Manages expression-variant image uploads, Ditto source registration, and image-to-video generation in one place.
Memory tabNotes kept across conversations
Shared / per-character long-term memory
Notes the AI carries across conversations. Shared applies to every character; per-character applies only to chats with that character.Store your name, preferences, and past events so it "remembers" you next time.
Advanced tabThe internals of speech recognition (STT) and speech (TTS)
System prompt
Sets instructions for the AI. Leave it empty for the default. Write the character's tone, role, and things to avoid.
Speech recognition (STT)
Language
Sets the base language for recognition and speech (Japanese / English / Chinese / multilingual auto). Multilingual mode auto-detects recognition language and lets VoxCPM2 synthesize multilingual text.
STT model
Chooses the engine that turns your voice into text. Whisper is general-purpose with a selectable size; Qwen3-ASR is high-accuracy (GPU recommended, heavy first load); Hiragana is a lightweight Japanese-only recognizer.Whisper size (tiny–large-v3): larger is more accurate but slower and uses more VRAM. On GPU aim for small–medium; on CPU, tiny–base.
Gemini raw audio input
When ON, sends the raw audio itself to Gemini in addition to the transcript, preserving pronunciation and intonation nuance.Only effective when a Gemini model is selected.
Speech (TTS)
TTS engine
Chooses how replies are turned into voice. WebSpeech is browser built-in and lightweight; VoiceVox is Japanese anime style; AivisSpeech is rich in emotion; VoxCPM2 supports multilingual Voice Design and cloning; Irodori lets you describe the voice in text.Irodori is controlled by writing the desired vibe in the "VoiceDesign caption." Roll 🎲 to land a voice you like and fix its "voice seed" so the timbre stays identical every time.
Speech speed
Sets how fast it speaks (0.5–2.0×).

Sample conversation

Pick a persona, talk into the mic, and the AI replies in voice. Avatars sync mouth movements when registered.

You
Be my English practice partner today — I have a job interview coming up.
Persona: Ryan (English tutor)
Ryan
Sure! Let's start with the classic question. Tell me about yourself — in 30 seconds.
Language: English / TTS: VoxCPM2
You
I'm a software engineer with five years of experience...
Ryan
Great opening. Try not to start every sentence with "I" — switch it up. Want me to give you a corrected version?
Pronunciation + grammar feedback

Supported voice languages

Multilingual voice input and output (VoxCPM2), focused on Japanese, English, and Chinese. Switch language per persona.

日本語 (JP)English (EN)中文 (ZH)粤語 (YUE)한국어 (KO)Español (ES)Français (FR)Deutsch (DE)Italiano (IT)Português (PT)Русский (RU)

Pairs with other features

Character Voice Chat gets stronger when combined with the rest of Kotonia — pick a partner from the public persona library, or jump into a packaged experience like Sassy AI English.

Try free voice chat now

One-minute signup, no credit card. 100 multilingual voice exchanges per day on the free tier.