Skip to content

Moonlight

Speech to text, builtfor the phone line.

Moonlight is our own recognition engine. Telephony audio goes in as it is, at 8 or 16 kHz, over WebSocket, gRPC or HTTP. Words come back while the person is still speaking, each one carrying the second it was said.

Contact Sales
connect.pyWebSocket
# stream telephony audio, get words back as they are spokenws = connect("wss://api.voicelab.ai/classify/asr", headers={"content-type": "audio/l16;rate=8000", # PCM, a-law, mu-law, FLAC"X-Voicelab-Conf-Name": "8000_en_US", # language + sample rate"X-Voicelab-Pid": "<your project>",})for chunk in audio:    ws.send(chunk)    print(ws.recv()){"status": "OK", "shift": "1", "words": ["the"]}{"status": "OK", "shift": "0", "words": ["there"]}{"status": "OK", "shift": "2", "words": ["is", "a"], "start": [1.84, 2.05]}
Live recognition - Inbound call - 8 kHz
  1. I am calling about the

    shift 4
  2. I am calling about the deliv

    shift 1
  3. I am calling about the delivery

    shift -1
  4. I am calling about the delivery

    final

Trusted by

How it works

Send audio. Getwords, with times.

One engine, three ways in. Stream over WebSocket or gRPC, or post files to the HTTP API, in the format your telephony already produces. One header decides which model listens.

  1. Send

    Stream over WebSocket or gRPC, or call the HTTP API. Telephony codecs go in as they are, and one header picks the configuration: the language and the sample rate together.

  2. Recognise

    Words come back while the person is still talking. Each partial result says how many of the previous words to replace, so what you show on screen is always the engine’s best current guess.

  3. Return

    JSON with the words and their start and end times, so you can jump to the second something was said, align it to your recording, or feed it straight into your own models.

Two things worth seeing

Live words, andthe time on each one.

Accuracy is the headline. These two are the parts that decide whether a transcript can carry a product on top of it.

Words while the sentence is still happening.

Partial results arrive within a breath of the sound, and each one says how far back to rewrite. Agent assist, live prompting and voicebots all need this, and a batch API cannot fake it.

  • Partial results with a shift, so text does not jump
  • A four byte marker ends the stream, nothing to poll
  • Example clients for all three protocols
Contact Sales

Every word carries the second it was said.

Start and end times per word are what turn a transcript into something you can build on: search that lands on the right moment, captions that sit under the picture, analytics tied back to the audio.

  • Word level start and end times
  • Plain JSON, nothing proprietary to parse
  • The same engine behind streaming and files
Contact Sales

Every way a buyer reaches you

What goes in,what comes out.

Configurations are named by language and sample rate, so a telephony line and a meeting recording are explicitly different models rather than the same one hoping for the best.

Contact Sales
Languages
  • Polish
  • English
  • German
  • Italian
  • Russian
Audio in
  • PCM 8 kHz
  • PCM 16 kHz
  • a-law
  • μ-law
  • FLAC
Protocols
  • WebSocket
  • gRPC
  • HTTP

Questions we get

Asked by everyengineering team.

If yours is not here, ask it before you integrate. We would rather answer it now than in a support ticket.

View all FAQs

What accuracy can we expect?

On your audio, measured against your reference, not on a public benchmark. We run thirty of your files in the first week and give you the word error rate next to the transcripts, so you can judge both. Anyone quoting a single percentage without hearing your recordings is quoting somebody else’s.

Does it work on 8 kHz telephony audio?

Yes. Moonlight accepts 8 and 16 kHz telephony audio, including PCM, a-law, μ-law and FLAC. Choose the configuration that pairs the language with the sample rate, then stream over WebSocket or gRPC, or send files through HTTP without resampling the source audio.

Which languages are available?

Moonlight supports Polish, English, German, Italian and Russian. The language and sample rate are selected explicitly in the configuration, so telephony audio and higher-bandwidth recordings use the appropriate model.

How long does an integration take?

It depends on the connection method and your production requirements, so we do not quote a fixed timeline before seeing the environment. We first confirm the audio format, language and sample-rate configuration, networking and acceptance test, then validate the integration and accuracy on your recordings.