CrossMeet / Guide

AI engine setup

Configure recognition, translation, speech, embeddings and other processing stages, checking hardware, connectivity and costs.

Updated

CrossMeet chooses engines by processing stage. Start with recognition and translation, then add speech, knowledge retrieval or document processing.

1. Identify the stages

Stage Purpose
ASR recognition Turn audio into original text
LLM translation Translate into the target language
TTS speech synthesis Speak confirmed text
Embeddings Generate vectors for knowledge entries and queries
Optional reranking Refine the order of retrieved candidates
OCR in document workflows Extract text from images or scanned content

Stages can use different configurations. Language, hardware and input support vary by model.

2. Choose local or API engines

Local engines require their runtimes and models and must pass hardware checks. Some require compatible GPUs; others support CPUs. Do not apply one engine's requirements to every feature.

API engines require a service address, model and credentials as applicable. Fields depend on the provider and interface. Use an account you are entitled to use and submit keys only to trusted endpoints.

Provider pricing, quotas and availability can change. This website does not maintain fixed third-party model prices or speed rankings.

3. Save and bind a configuration

Add or edit an engine in settings, check connectivity and model availability, then bind it to the intended task.

Adding a configuration does not make every workflow use it. Meeting translation, document translation and knowledge embeddings can have different bindings. When troubleshooting, inspect the active binding rather than only the list of saved models.

4. Run a short test

  1. Use a short, non-sensitive audio or text sample.
  2. Check ASR output and confirm the input device and language.
  3. Check translation, target language and context.
  4. If needed, test the TTS voice and output device.
  5. Change one setting at a time and observe the result.

Noise, segmentation, models, hardware and networks affect the experience. No single parameter set suits every meeting.

5. Customize prompts

Prompt settings list the variables available for each template. Preserve required placeholders and test changes on a short sample.

For example, a translation template supporting source and target languages can include:

Translate from {src_lang} to {tgt_lang}.
Return only the translation.

{src_lang} is the source language and {tgt_lang} the target. Other templates can have different variables. Do not copy variables that the selected template does not list.

Manage preferred terms in the glossary and business Q&A in the knowledge base.

6. Team LAN engine sharing

Team can share compatible translation, recognition, speech, OCR, embedding and reranking services from a configured LAN host. Actual availability depends on its engine types, models and configuration.

Choose the sharing scope and manage access tokens on the host. Other devices connect with the corresponding address and token. Test permissions and capacity on an approved network. Capacity depends on the host; no fixed user count per machine is promised.

LAN engine sharing does not automatically distribute members' glossaries, phrases or knowledge entries. Use each module's supported imports and exports to manage personal material.

7. Data and costs

Check local, API and LAN boundaries at every stage. Cloud embeddings, translation, speech and reranking send relevant text or audio to their providers. See privacy.

Free, Pro and Team define software access. Bring-your-own API usage is billed separately. Subtitle Bridge official cloud also has its own points subscription; it is not a universal free allowance for desktop engines.

If something fails

Check the active binding, model name, credentials, provider quota, network and local model status. Follow troubleshooting for more detailed checks.

Your next conversation

Start with a clearer understanding.

Explore Free, then add the tools your work needs.