This is the backbone every voice call runs through: speech-to-text turns audio into a transcript, then the NLU engine classifies intent and pulls out entities. Test it live below.
Try it
Speak into your microphone or type a sentence — both run through the same pipeline.