Capabilities · 13 modules

Engineered for serious cross-language conversations.

From understanding, to speaking, to professional correction, to team sharing — full pipeline covered.

~100ms
ASR-to-screen latency
4
ASR engines
30+
Languages
Parallel LLM streams
01 · Real-time

Bidirectional simul-interpret

Two audio streams (your mic + system audio) drive ASR + translation in parallel. Left/right panels show your side and theirs separately. Turn-taking is instant.

  • Source text in 100ms, translation streams in parallel
  • Mic + system audio independent ASR pipelines
  • Final segment = LLM-corrected source + translation pair
EN · Source
"How's the project going?"
"We need the deck by Friday."
ZH · Translation
项目进展如何?
我们需要周五前完成幻灯片。
02 · Hybrid

Local + cloud hybrid ASR

Whisper-faster (local GPU/CPU), Whisper API, Aliyun ASR, Qwen3-ASR — choose per scenario. Local first for privacy; cloud for max precision.

LOCAL · GPU
Whisper-faster
CTranslate2 · 4× speed
CLOUD · API
Whisper-1
OpenAI compatible
CLOUD · CN
Aliyun DashScope
native realtime ASR
CLOUD · CN
Qwen3-ASR
low-latency streaming
03 · Languages

30+ languages, any pair

Chinese, English, Japanese, Korean, German, French, Spanish, Russian, Portuguese, Arabic, Hindi, Vietnamese, Thai, Turkish, Polish — and more. Auto-detect or manual locale lock.

中文
English
日本語
한국어
Deutsch
Français
Español
Русский
Português
العربية
हिन्दी
+ 19
04 · Accuracy

Glossary management

Custom brand names, person names, technical terms — three categories. Group-based activation per meeting. Injected into LLM prompt and validated post-hoc.

BEFORE → AFTER
"You should try Cursor for coding"
→ 你应该试试光标来编程
+ Glossary: Cursor (product name)
"You should try Cursor for coding"
→ 你应该试试 Cursor 来编程
05 · Context

Knowledge base RAG

Upload company materials, product docs, customer briefs. Vector retrieval injects context into the translation prompt. LLM understands your business.

WORKFLOW
Upload PDF / DOCX → vector index
ASR text → similarity search
Top-k snippets injected to prompt
LLM translates with full context
06 · Speaking

Virtual microphone output

TTS output routes to a virtual microphone. Zoom / Teams / Discord receive translated native voice. No copy-paste, no software switching.

SIGNAL FLOW
  1. Your mic muted (no audio sent)
  2. You type or speak in native
  3. LLM translates & refines
  4. TTS speaks into the virtual mic
  5. Zoom / Teams / Discord hears native voice
07 · Precision

Compose Bar (simul-interpret)

Type in your native language, see translation preview, manually correct, then send only when 100% correct. Zero translation accidents.

WORKFLOW
  1. 1 Type in native language
  2. 2 Live translation preview as you type
  3. 3 Manual edit if needed
  4. 4 Send: TTS + virtual mic
08 · Fluency

Quick phrases

Pre-set common phrases (greetings, small talk, business asks). Send any phrase with one click — text + TTS go out together via the virtual mic.

EXAMPLE PHRASES
Could you repeat that, please? ⌘1
Let me check and get back to you. ⌘2
Sounds good — let's proceed. ⌘3
09 · Recap

History & export

Every session is recorded automatically. Search, re-translate, summary generation, Markdown / PDF / DOCX export.

Q3 client meeting 2026-04-15 · 47min
Acme onboarding 2026-04-12 · 23min
Daily standup 2026-04-10 · 8min
.md .pdf .docx
10 · Team

LAN API sharing

Expose local engines (ASR / TTS / LLM) as OpenAI-compatible APIs to other devices on your LAN. Team shares one GPU machine — saves cost.

curl example
curl http://192.168.1.42:18000/v1/chat/completions \
  -H "Authorization: Bearer $LAN_TOKEN" \
  -d '{"model":"local-llama","messages":[...]}'
11 · Setup

Cloud platform presets

OpenAI, Anthropic, Aliyun, DeepSeek, SiliconFlow — paste your API key, the engine list auto-populates. Switch providers in one click; no JSON wrangling.

SUPPORTED PRESETS
OpenAI
Anthropic Claude
Aliyun DashScope
DeepSeek
SiliconFlow
Ollama (local)
Gemini
+ custom
12 · Privacy

Local-first by design

Whisper-faster + local TTS run entirely on your GPU/CPU. Sensitive meetings can be 100% offline. Cloud APIs are strictly opt-in — and even then, requests go from your machine directly to the provider, never proxied through CrossMeet servers.

DATA FLOW
  • Audio never leaves your machine in local mode
  • Cloud requests go direct to provider; we never proxy
  • License + telemetry are minimal; logs auto-purge in 1h
  • Hardware gating: no NVIDIA? local features auto-disable
13 · Learning

CrossLearn — learn languages by watching

Embedded video player with smart subtitles. Pause on hover, double-click any word to save into your vocab. AI-generated dictionary cards explain context, etymology, and usage — built for active learning, not passive scrolling.

  • YouTube + Netflix + Bilibili (via Subtitle Bridge browser extension)
  • Hover-pause + word lookup with dictionary card
  • Double-click to save into your personal vocab notebook
HOW IT WORKS
  1. 1 Open a video page in browser, extension pushes subtitles to CrossMeet
  2. 2 Subtitles render with vocab highlighting in real time
  3. 3 Hover a word to pause + see dictionary card
  4. 4 Double-click → save to vocab notebook for review

How CrossMeet compares

Feature CrossMeet Otter.ai Krisp Google Translate
Real-time bilingual
Local processing (privacy)partial
Virtual microphone output
Knowledge base RAG
Glossary managementlimited
30+ languagesEN onlyEN only
One-time pricing
Get started

Ready to make every conversation feel native?

NO CREDIT CARD CANCEL ANYTIME LOCAL-FIRST
~100ms
End-to-end latency
30+
Languages
4
ASR engines
WIN 10/11
Native platform