# Creating a Knowledge Base
Source: https://docs.gnani.ai/A01_KB
### **Introduction & Context**
A robust knowledge base is the foundation of your AI agent, it contains the information your agent uses to generate accurate responses. Think of it as the agent’s **brain**, filled with knowledge that enables it to answer questions effectively.
*Your agent’s brain needs fuel. Let’s feed it!*
***
### **Lesson 1: Uploading a Document (PDF/DOC)**
* **What It Is:** Upload documents that contain essential information such as FAQs, manuals, company policies, and SOPs to train your AI agent.
* **How It Works:**
1. Navigate to **Knowledge Base → Add Knowledge Base.**
2. Click in the **Knowledge Base Files** section and select your PDF, DOC, DOCX, XLS, XLSX, CSV or TXT file(s).
3. You can also drag and drop your files.
4. Click on **Next** and give you Knowledge Base a name.
5. Click on **Start Training** to process the document.
6. The platform analyzes the content and makes it accessible for the genAI agent to use.
* **Use Cases:** Customer FAQs, internal policy documents, product manuals.
* **💡 Pro Tips:**
* Use clear headings and structured formatting. Your genAI agent learns best from well-organized documents.
* Need to update your knowledge base? Go to the **Knowledge Base** section and click on the delete or upload button
***
### **Lesson 2: Importing a Knowledge Base from a URL**
* **What It Is:** Extract content from a webpage to train your AI agent.
* **How It Works:**
1. Navigate to **Agent Related → Knowledge Base → Add Knowledge Base.**
2. Click on **Add a URL.**
3. Paste a valid website URL and click on **Next.**
4. The platform will scrape text content from the page and follow links (up to Level 1).
5. Select specific sub-URLs to include, then click **Next** and enter a name for your Knowledge Base.
6. Click **Start Training.**
* **Why Use This?** Perfect for leveraging online help centers, publicly available documents, or product pages without manually copying content.
* **Use Cases:** Company blogs, support portals, product pages, pricing pages.
* **💡 Pro Tips:**
* Ensure the webpage is publicly accessible—password-protected sites won’t work.
* If the page contains excessive unrelated content, consider uploading a document instead for better accuracy.
***
🎉 *Great job! Your genAI agent’s brain is now loaded with knowledge!*
✅ **Next Steps:**
* Click **Test Knowledge** to see how your AI responds using the uploaded content.
* Continue refining your knowledge base by adding, updating, or removing documents and URLs as needed
# Creating Your first Agent
Source: https://docs.gnani.ai/A02_Agent
Time to roll
### **Introduction & Context**
Creating an agent is where the magic happens! This is where you create your GenAI assistant that interacts with users, making conversations based on the knowledge base and prompt that you've added.
### **Lesson 1: Starting from Scratch or Using a Template**
* **From Scratch:**
* **When to Use:** When you want full control and customization.
* **How:**
* Go to **Manage Agents → Create Agent → Create from Scratch**.
* Name your agent, give a description for your understanding and click on **Create**.
* **Using a Template:**
* **When to Use:** When you need a quick start, use templates as they include pre-written system prompts.
* **How:**
* Go to **Manage Agents → Create Agent.**
* Choose a template that best fits your needs and click on Proceed. You can customize your agent later.
* Name your agent, give a description for your understanding.
* Link a knowledge base and click on **Create**.
### **Lesson 2: Writing a Great System Prompt**
* **What It Is:** A system prompt defines your agent's role, behavior, and objectives.
* **Tips for a Great System Prompt:**
* **Set Clear Objectives:** Provide specific details about what you want your agent to accomplish.
* **Give Context:** Provide some background information about the role, audience and purpose.
* **Set Expectations:** Explain the tone, style, and level of formality.
* **Include Examples:** Provide scenarios to guide responses.
### **Lesson 3: Setting the Model**
* **Greeting Message:** Set the exact message your agent says at the start of the call. You can add dynamic variables to personlize your greeting message.
* **Ending Message:** Define the closing statement your agent delivers before ending the call.
* **Provider & Model:** Choose the LLM that powers your interactions.
* **Link Knowledge Base:** Connect your previously created knowledge base so your agent has information to draw upon.
* **Temperature:** This controls creativity. A higher value means more creative (but sometimes less precise) responses, whereas a lower value means conservative responses (sticking strictly to the knowledge base).
* **Max Tokens:** Limits the length of user's responses that the agent can process in one input. A token is usually 3/4 of an English word or 3-4 characters.
### **Lesson 4: Customizing your Agent**
Fine-tuning your agent's settings ensures it not only understands but also reflects your brand's personality. This section is where you polish the details. Head over to the **Customize** tab in Manage Agent section.
**Agent Details:**
* **Language:** Select the language(s) your agent communicates in. Choose a primary language if you're selecting multiple languages.
* **Region & Time Zone:** Select the relevant Region and the Time Zone for your agent. This will set the time and will be used further (when integrations are used).
* **Description:** Specify the way your agent talks with the user.
# Testing your Agent
Source: https://docs.gnani.ai/A03_Testing
**Testing via Chat:**
* Click on **Test → Chat Window → Start Testing \>** to use the built-in chat interface to test responses.
**Testing Voice Interactions:**
* Click on **Test → Web-based (Voice) → Start Testing \>** to use the web-based voice test to simulate voice interactions.
* To share the web-based voice testing tool, go to **Test → Web-based (Voice) → Generate Sharable Link**. The link works for 5 minutes, perfect for giving teammates or clients quick access without login requirements.
**Testing a real-world call:**
* Click on **Test** **→** **Trigger Agent Call** and select your whitelisted number from the dropdown.
* Click **Start Testing** and wait for your call.
*Tip:* Keep your phone handy during testing, you’re about to experience your agent in action!
***
*Pro Tip:* Experiment with different temperatures, system prompts, transcribers and text-to-speech to see what best suits your use case. The right balance can make your agent more engaging, effective and tailored.
# Whitelisting Numbers
Source: https://docs.gnani.ai/A04_Whitelisting
Simulate a real call
### What Are Whitelisted Numbers?
Whitelisted numbers allow you to test your agent’s calling feature. By adding your own phone number, you can receive test calls and interact with your agent in real time, ensuring everything works as expected.
### Why Use Whitelisted Numbers?
* **Real-World Testing:** Experience firsthand how your agent interacts over a phone call.
* **Seamless Integration:** Once added, your number appears in the test call dropdown, allowing you to trigger calls for any agent effortlessly.
### How to Add a Whitelisted Number
1. **Access the Whitelisted Page:**
* In the sidebar, click on **Phone Numbers** and then select the **Whitelisted** page.
2. **Add Your Number:**
* Click on **Add Number** in the top right.
* Enter your phone number and assign it a friendly name for easy identification.
* Click **Next**.
3. **Verification:**
* An OTP (One-Time Password) will be sent to your phone.
* Enter the OTP in the provided field and click **Verify**.
4. **Test Your Agent:**
* Navigate to **Manage Agents** → select an agent → click on **Test**.
* Click **Trigger Agent Call** and select your whitelisted number from the dropdown.
* Click **Start Testing** and wait for your call.
*Tip:* Keep your phone handy during testing, you’re about to experience your agent in action!
# DTMF Collection
Source: https://docs.gnani.ai/B01_DTMF
Let your voice agents listen, capture, and respond to **keypad inputs** from users — essential for collecting sensitive inputs like phone numbers and PIN codes securely and efficiently.
**DTMF (Dual-Tone Multi-Frequency)** support ensures your agents can handle numeric inputs with precision, whether it's authenticating users or routing calls based on input.
## What is DTMF and Why Enable It?
DTMF is a system that sends numeric input over phone lines — like when you enter a PIN, press "1 for support", or key in an account number. With DTMF enabled, your voice agents can collect:
* **Phone Numbers**
* **PIN Codes**
* **Account Numbers or OTPs**
**Why it matters:**
* **Secure Input:** Collect sensitive information without relying on voice transcription.
* **Better UX:** Give users a familiar way to respond via dial pad.
* **Essential for IVR:** Key component of any interactive phone workflow.
* **Improves Accuracy:** Ideal for situations with background noise or low-quality audio, where speech recognition may struggle.
## How to Enable DTMF Support
You’ll find a toggle under **Agent Settings → Customize → DTMF Collection**.
> ⚠️ **Important:** Simply turning on the DTMF toggle does **not** activate the full flow.\
> To make it work, you must also update your **system prompt** using the correct DTMF signals as shown below.
## DTMF Collection Template
Update your system prompt with the following format when collecting inputs. The signal format must be **precisely** followed.
### Collecting Phone Number
When your agent asks for a phone number:
* Append this signal: `| DTMF1010`
* Ask the user clearly to provide input **after the beep**
* Input must be a 10-digit number (e.g., `8197800293`)
**System Prompt Example:**
> "Can you please provide your phone number after the beep tone? | DTMF1010"
**If the input is invalid:**
> "That doesn't seem like a valid phone number. Please enter a 10-digit phone number after the beep. | DTMF1010"
### Collecting Pin Code
When your agent asks for a PIN:
* Append this signal: `| DTMF0610`
* Input must be exactly 6 digits (e.g., `560033`)
**System Prompt Example:**
> "Can you please provide your pin code after the beep tone? | DTMF0610"
**If the input is invalid:**
> "That doesn’t seem like a valid pin code. Please enter a 6-digit pin after the beep. | DTMF0610"
## How DTMF Signals Work
DTMF signals follow the format:\
`| DTMF[XXYY]`\
Where:
* `XX` = Expected number of digits
* `YY` = Time (in seconds) allowed for user input
For example:
* `DTMF1010` → Expect **10 digits**, allow **10 seconds**
* `DTMF0610` → Expect **6 digits**, allow **10 seconds**
> ⚠️ **Tip:** Always set the `YY` duration based on the TTS utterance length.\
> If the message is long but the time is short, the input might **timeout before the beep.**
## 💡 Bonus Tips & Best Practices
| Tip | Why It Helps |
| ---------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| **Match time to utterance length** | If the `YY` value (input time) is too short for your message, users won't be able to respond in time. |
| **Don’t strip or reformat DTMF signals** | They must remain at the end of the prompt **exactly** as written. |
| **Handle invalid input loops** | Guide the user to retry if their input doesn’t match the expected format. |
| **Use clear instructions** | Mention “after the beep” to guide user expectations and ensure input is captured. |
| **Enable barge-in (optional)** | If you're using TTS, enabling barge-in allows users to enter input **without waiting** for the entire prompt to finish. |
| **Validate inputs post-call** | You can log or validate the collected digits after the call for reporting or verification workflows. |
| **Use in noisy environments** | DTMF ensures input collection even when speech transcription is unreliable due to background noise. |
By enabling and correctly using DTMF, you open up powerful, secure, and user-friendly workflows for your voice agents — from authentication to onboarding and beyond.
Let your agents hear more than just voices. Let them listen to actions.
# Advanced ASR Settings
Source: https://docs.gnani.ai/B02_Advanced_ASR
Control how your voice agent listens, detects, and processes user speech in real-time conversations.
### **Overview**
Advanced ASR (Automatic Speech Recognition) settings control how your AI voice agent listens and responds to spoken input.\
They define listening limits, silence detection, interruption behavior, and background noise filtering for optimal call experience.
### Location in Platform
**Manage Agent → Customization → Advanced ASR Settings**
> **Note:** These settings apply only to **voice channels**. They do not affect chat or text-based agents.
### **Available Settings**
| **Setting** | **Description** |
| :-------------------------------------- | :----------------------------------------------------------------------------------- |
| **Max Speech Duration** | Maximum duration (in seconds) the agent will listen to a single user input |
| **Initial Silence Timeout** | Time to wait for the user to start speaking before cancelling |
| **End Silence Timeout** | Time to wait after user stops speaking before finalizing the input |
| **Speech Segmentation Silence Timeout** | Mid-speech silence duration used to split long speech into segments |
| **Allow Interruptions** | Lets the user speak over the agent; the agent stops speaking and listens immediately |
| **Interrupt Initial Message** | Allows interruptions during the agent’s very first message in a call or interaction |
| **Background Noise Filtering** | Adjusts sensitivity to background sounds to reduce false interruptions |
## **How Each Setting Works**
### **1. Max Speech Duration**
**Description:** Specifies how long the agent will listen to user input in one stretch before automatically stopping.
**Use Case:** Prevents prolonged listening due to background noise or over-talking. Ensures the agent maintains a responsive, controlled interaction.
### **2. Initial Silence Timeout**
**Description:** Defines how long the agent will wait for the user to begin speaking at the start of a turn. If the user remains silent beyond this threshold, input is cancelled or retried.
**Use Case:** Useful when users are unsure, distracted, or take time to process the prompt. Prevents the system from hanging indefinitely.
### **3. End Silence Timeout**
**Description:** Sets the duration of silence after the user stops speaking, after which the agent considers the input complete and proceeds.
**Use Case:** Allows for natural pauses while still keeping the interaction smooth. Essential for avoiding premature cutoff.
### **4. Speech Segmentation Silence Timeout**
**Description:** When users speak in long sentences or paragraphs, this setting helps break the input into segments based on silence detection. Particularly useful for streaming ASR or multi-sentence inputs.
**Use Case:** Improves comprehension and reduces memory load on the model by processing inputs in manageable chunks.
### **5. Allow Interruptions**
**Description:** When enabled, the agent will stop speaking and immediately listen when the user starts talking.\
**Use Case:** Creates a more natural, back-and-forth conversation where users can cut in.
### 6. Interrupt Initial Message
**Description:** When enabled (and **Allow Interruptions** is ON), the agent will also allow interruptions during its very first message.\
**Use Case:** Lets impatient users respond immediately, even during the greeting.
### 7. Background Noise Filtering
**Description:** Controls how sensitive the interruption feature is to background sounds.\
Low sensitivity (closer to 20): More likely to trigger on quiet sounds, including unwanted noise.\
High sensitivity (closer to 100): Better at ignoring noise like traffic or barking but may miss very soft speech.\
**Use Case:** Reduce false triggers while balancing responsiveness.
### Recommendations
* Use **shorter timeouts** for transactional bots (e.g., booking, verification).
* Use **longer timeouts** for support scenarios, complex discussions, or with elderly users.
* Enable **Segmentation** when expecting detailed or multi-part answers.
* Enable **Allow Interruptions** for more natural, human-like interactions.
* Adjust **Background Noise Filtering** based on the expected environment.
# Voicemail Detection
Source: https://docs.gnani.ai/B03_Voicemail
Voicemails are a part of outbound calling but missed opportunities and awkward message drops don’t have to be.
The **Voicemail Detection** feature helps your voice agent smartly identify when it’s reached a voicemail inbox and respond accordingly, avoiding awkward or wasted greetings, and optionally triggering automated follow-ups like SMS or email.
This is especially useful in outbound call flows where pickup detection and timing are critical, such as **sales outreach**, **appointment reminders**, or **lead qualification**.
🧭 **Path**: `Manage Agent → Customize tab → Voicemail Detection`
## Why Voicemail Detection Matters
* Improve agent experience by avoiding cutoff intros or beeps.
* Save on costs by skipping unnecessary conversation attempts.
* Boost user trust with professional, voicemail-appropriate responses.
* Automate engagement by triggering SMS, emails, or workflows after detection.
## What It Does
When an outbound call is answered, your agent uses real-time audio cues to detect whether the recipient is a person or a voicemail system. Based on your configuration, it will either:
* Retry voicemail detection,
* Wait for the right moment to respond,
* Speak a custom voicemail message,
* Trigger automated follow-ups.
## What You Can Configure
### 1. Max Detection Attempts
* **What it is**: Number of times the system retries to determine if the call went to voicemail.
* **Why it matters**: Helps balance detection accuracy vs. speed and cost.
* **Range**: 1 to 5
* **Example**: Setting this to 3 means the system will attempt detection up to 3 times before deciding.
*Tip: Increase for better accuracy; decrease for faster calls and lower compute cost.*
### 2. Voicemail Playback Delay
* **What it is**: Number of seconds the agent waits after voicemail is detected before speaking.
* **Why it matters**: Prevents the agent from being cut off by long greetings or carrier beeps.
* **Range**: 1 to 10 seconds
* **Example**: A delay of 4 seconds gives enough time for most voicemail intros to end before the message plays.
*Why delay?*\
Voicemail greetings can include:\
“Hi, this is John. I’m not available right now…” **\[beep]**\
If your agent starts too early, its message may be lost.
### 3. Voicemail Response
* **What it is**: The message your agent will speak once voicemail is detected. You can also use dynamic variables in the message by using double curly braces like `{{customer_name}}`
* **Limit**: 300 characters max
* **Example**:\
`"Hi, we tried reaching you but missed you. Please call us back or reply to this message."`
*Make it short, polite, and clear. Don’t include long messages. The goal is clarity, not conversation.*
### 4. Post Detection Actions
* **What it is**: Optional actions your agent can take after leaving a voicemail.
* **Examples**:
* Send a follow-up SMS (e.g., "Just left you a voicemail. Text us back!")
* Send an email summary to the contact owner
* Trigger a custom API
*Use this to automate re-engagement or log activity in CRMs or backend systems.*
## Bonus Tips & Best Practices
| Tip | Why It Helps |
| --------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Keep initial greeting short | If the bot’s first message is too long, it might overlap with the phone’s voicemail greeting, making it impossible to capture user audio or transcripts. Keep it concise, especially in the first 5–10 seconds. |
| Enable Barge-In Detection | Turn on Barge to let your agent listen while speaking. This improves voicemail detection accuracy, especially when voicemail systems start mid-sentence or play long intros. |
| Personalize the voicemail message | A natural, human-sounding message improves callback rates. Add a name or reason for the call. |
## Example Scenario
1. Agent makes an outbound call.
2. No human detected → system retries 2 more times.
3. Detection succeeds on attempt 3.
4. Agent waits 5 seconds → speaks voicemail message.
5. Triggers CRM update.
## When Should You Use It?
* When your outbound calls frequently hit voicemail.
* If you want to leave clean, professional messages without being cut off.
* When you want to automate follow-up actions after voicemails.
* To improve response rates by sending SMS/email right after missed calls.
## What This Feature Does Not Do (Yet)
* It doesn’t guarantee 100% voicemail detection.
* It won’t auto-detect specific voicemail messages (e.g., “the mailbox is full”).
## Best Practices
| Goal | Setting Recommendation |
| ----------------------- | -------------------------------------- |
| Max accuracy | Max Attempts: 5, Delay: 4–5s |
| Faster call completions | Max Attempts: 1–2, Delay: 2–3s |
| Better callback rate | Always add Post Detection SMS or Email |
| Avoid being cutoff | Use a delay of at least 3s |
## Troubleshooting
| Issue | What to check |
| ---------------------- | ----------------------------------------------------- |
| Agent speaks too early | Increase Voicemail Playback Delay |
| Message not played | Check if detection attempt limit was too low |
| No follow-up actions | Ensure Post Detection Actions are configured properly |
# Dynamic Variables & Dynamic Messages
Source: https://docs.gnani.ai/B04_Dynamic_Variables
Personalize your agent's responses
The **Dynamic Variables** feature in Gnani Agents allows you to insert placeholders in your bot’s **Greeting Message** or **System Prompt**, which get replaced with real-time data during a call.\
These variables are populated from an API configured in the **Dynamic Messages** settings.
***
### **Dynamic Messages Overview**
When you turn on **Dynamic Messages** in the agent settings, you will be asked to configure:
* **Method** — `GET` or `POST` (required)
* **API URL** — endpoint that returns the dynamic message and user context (required)
* **Headers** — optional key–value pairs (if applicable)
This API request will contain:
1. A **unique variable** (for example, the customer’s phone number).
2. Return:
* **User context** — any key–value data needed for the call (e.g., `due_amount`, `customer_type`, `due_date`) in JSON format.
* **Constructed greeting message** — the exact greeting to be used for that specific call.
⚠ **Important:**
* The Greeting Message configured in the agent will be **overridden** by the message returned from your API when Dynamic Messages is on.
* The API call has a **10-second timeout**. If it doesn’t respond in time, the call will fail.
* Currently, Gnani Agents does **not** have a built-in campaign manager. This feature is available only by contacting us to enable it for your deployed bot. We’re working to make it self-service soon.
* The **unique variable** can be something other than the phone number.
***
### **Using Dynamic Variables in the System Prompt**
Dynamic variables can also be used inside the **System Prompt**, but:
1. **Dynamic Messages** must be turned on.
2. In the **Customize** section, enable **Pre-Call Variables**.
3. Add the required variables you plan to use in your System Prompt.
**Testing behavior:**
* During test calls, you’ll be prompted to manually enter the test variable values.
* In production, these will be **auto-filled** from your Dynamic Messages API response.
***
### **Where Dynamic Variables Can Be Used**
✅ **Supported:**
* Greeting Message (overridden by API response)
* System Prompt
❌ **Not Supported:**
* Ending Message
***
### **API Specification**
**Request Body**
When the API is called, Gnani Agents sends:
```json theme={null}
{
"conversation_id": "abc123",
"mobile": "01234567890"
}
```
**Expected Response**
```json theme={null}
{
"additional_info": {
"inya_data": {
"text": "Hi John Doe, your payment of ₹2500 is due on 2025-08-20.",
"user_context": {
"phone_number": "+911234567890",
"name": "John Doe",
"due_amount": "2500",
"due_date": "2025-08-20",
"customer_type": "Premium"
}
}
}
}
```
* `text` → Greeting message to be used for the call (mandatory).
* `user_context` → Key–value pairs for use in the System Prompt (mandatory; user\_context can be an empty dictionary if no dynamic variables are used).
***
### Sample API in Python
```python theme={null}
@router.post("/initial_message")
async def bot_initial_message(request: Request):
body = await request.json()
mobile = body.get("mobile")
# Example: fetch data from your CRM
user_data = crm_query(mobile)
response_data = {
"additional_info": {
"inya_data": {
"text": "Hi {user_data.get('name')}, this is a test message",
"user_context": {
"phone_number": mobile,
"name": user_data.get("name"),
"age": user_data.get("age")
}
}
}
}
# Status code 200 → call will be triggered
# Status code 400 → call will NOT be triggered
return JSONResponse(content=response_data, status_code=200)
```
> This API can be built in any language. The above is a **Python** example.
***
## Best Practices
* Keep API latency low to avoid hitting the 10-second timeout.
* Return only the variables you need to keep payloads small.
* Always validate that your API returns both `text` and `user_context`.
* Use consistent naming for variables so they can be reliably inserted into prompts.
# Writing a Disposition Prompt
Source: https://docs.gnani.ai/C01_Disposition
Customize the post-call analytics
Given below is the structure of the disposition prompt. Follow the structure to craft your own call-transcript extraction prompts:
### 1. Define the Role and Context
* **What does it do?**\
Specify who the LLM is and the domain it operates in (e.g., Call Center Operations Analyst in Debt Collection).
* **Why it matters:**\
Sets expectations and tailors the model's responses.
### 2. Clarify the Task
* **What does it do?**\
Write a brief `"## Task"` section that explains exactly what the model should accomplish (e.g., read transcript, extract fields).
* **Why it matters:**\
Keeps the model focused on the end goal.
### 3. Emphasize Output Requirements
* **What does it do?**\
Under `"## IMPORTANT"`, stress the format (JSON) and forbid extra fields or custom codes.
* **Why it matters:**\
Ensures structured, machine-readable output.
### 4. Provide Business and Language Context
* **What does it do?**\
Use `"## Context Understanding"`, `"## Business Context"` and `"## Language Processing Guidelines"` sections to share domain rules, use cases, and allowed languages.
* **Why it matters:**\
Guides the model on tone, terminology, and multilingual handling.
### 5. List Field Definitions
* **What does it do?**\
Create a JSON snippet showing each key with an empty value.
* **Why it matters:**\
Shows exactly which fields to populate and in what structure.
### 6. Detail Allowed Values
* **What does it do?**\
Under `"### Allowed Values & Definitions"`, enumerate each field's valid codes, descriptions, and criteria in priority order for STAGE\_CODE.
* **Why it matters:**\
Prevents misclassification and enforces consistency.
### 7. Add Critical Analysis Instructions
* **What does it do?**\
Numbered guidelines on transcript reading, stage determination, priority rules, and callback handling.
* **Why it matters:**\
Helps users correctly apply the template and avoid common mistakes.
### 8. Show Output Format Example (Optional)
* **What does it do?**\
Provide a sample JSON response matching the field definitions.
* **Why it matters:**\
Offers a quick reference for the expected output.
### 9. Insert the Transcript Placeholder
* **What does it do?**\
End with a `"## Transcription"` header where the actual call log goes.
* **Why it matters:**\
Clearly demarcates where to paste the raw data.
***
## Tips
* Copy the template and fill in the `{{placeholders}}` with your specific details.
* Use simple, clear language when defining the call purpose and rules.
* Don't remove or rename any keys in the JSON snippet, they must match exactly.
### Example Prompt Template
````mdx theme={null}
## Role
You are a skilled Call Center Operations Analyst specializing in {{Industry}}
operations. You will be given call logs that contain detailed conversation
transcripts between an Agent and a User. The call transcripts could be in
{SUPPORTED_LANGUAGES} or mixed language.
## IMPORTANT
Provide your response strictly in **JSON format** following the specifications
below. Do not introduce any additional fields or custom stage codes beyond those
defined.
## Task
Read the entire conversation transcript carefully and extract the required
information according to the rules. Deliver a structured JSON response containing
only the specified keys.
### Context Understanding
- **Call Purpose**: {{Call Purpose}}
- **Participant Roles**: Agent and User
### Business Context
- **Industry**: {{Industry}}
- **Use Case**: {{Use Case}}
- **Business Rules (Optional)**:
- 1. {{Rule 1}}
- 2. {{Rule 2}}
### Language Processing Guidelines
- **Primary Language**: {{Primary language}}
- **Secondary Languages (Optional)**: {{Secondary language (if applicable)}}
## Specifics
### Field Definitions
```json
{
"STAGE_CODE": ""
}
```
### Allowed Values & Definitions
#### STAGE\\\_CODE Values:
```json
[
{
"code": "ESCALATED_TO_AGENT",
"description": "Transferred to human agent for complex issues",
"criteria": "User requests escalation or dispute is detected."
},
{
"code": "AGREES_FOR_CALLBACK",
"description": "User agrees to a callback",
"criteria": "User explicitly agrees to be called back."
},
{
"code": "DISAGREES_FOR_CALLBACK",
"description": "User declines a callback",
"criteria": "User explicitly declines being called back."
},
{
"code": "BUSY",
"description": "User indicates they are busy",
"criteria": "User says they cannot talk now."
},
{
"code": "FAQ_HANDLED",
"description": "User question answered without escalation",
"criteria": "User asks a routine question and receives answer."
},
{
"code": "NO_INPUT",
"description": "No user response after prompt",
"criteria": "Silence or no intelligible input."
},
{
"code": "INVALID_INPUT",
"description": "Unrecognized or unclear response",
"criteria": "User input is garbled or irrelevant."
},
{
"code": "WRONG_NUMBER",
"description": "Wrong number reached",
"criteria": "User indicates wrong number."
},
{
"code": "DND",
"description": "Do-not-disturb request",
"criteria": "User requests not to be contacted again."
}
]
```
## Critical Analysis Instructions
1. Read the **ENTIRE transcript** before extracting. Don't jump to conclusions
2. **Look for explicit user responses** - Don't assume agreement from silence
## Important Guidelines
### Data Extraction Rules
- **Key Restriction**: Extract only the keys specified in the field definitions.
Do not add any additional keys or fields to the output.
- **Output Compliance**: Follow the exact JSON format specified. The output must
contain only the fields defined in the field definitions section.
## Output Format
Provide response in JSON format.
### Standard Response Structure - json
```
{
"STAGE_CODE": "{EXTRACTED_VALUE}"
}
```
## Transcription
\
````
***
# Language Switch Prompt
Source: https://docs.gnani.ai/C02_Language_Switch
Configure how your agent detects and responds to user requests for changing languages during a conversation.
### **Overview**
The Language Switch Prompt enables your AI agent to dynamically identify and respond when a user explicitly or implicitly requests a switch in language. This functionality is critical for agents that support multilingual conversations, especially over voice where real-time language adaptation enhances usability and reach.
### **Purpose**
When multilingual agents are enabled, this prompt instructs the system to:
1. Detect whether the user is requesting a language change.
2. If detected, identify the specific target language.
3. Return the correct output so that downstream logic can route accordingly.
### **Where It Lives**
**Location:**\
Manage Agents → Overview → Language Switch Prompt
Whenever you **add or remove a language**, the system prompts you to update this field to ensure alignment between what the agent supports and what it can detect.
A pre-filled **reference prompt** is provided. You are expected to customize it based on the languages added to your agent and their corresponding trigger phrases.
### **Why It’s Important**
Language switching relies on both **system configuration** and **prompt logic**. If a language is added to the agent configuration but not reflected in the prompt (or vice versa), switching will fail at runtime. The prompt becomes the source of truth for how language detection is handled across STT engine outputs.
### **Prompt Template (To Customize)**
Below is the reference structure you should adapt based on your agent’s supported languages and expected phrases.
```markdown theme={null}
## Role:
- You are a specialized intent detection agent designed to analyze outputs from multiple Speech-to-Text (STT) engines and identify if the user is explicitly requesting to switch or change languages.
## Objective:
- Determine if the user is requesting a language switch by examining transcriptions from multiple STT engines.
- If a language switch intent is detected, identify the specific language the user wants to switch to.
## Input Format:
- S1: Engine 1 Output (can be in any script/language)
- S2: Engine 2 Output (can be in any script/language)
- S3: Engine 3 Output (can be in any script/language)
## Output Format:
- If a language switch intent is detected, output ONLY the target language name.
- If no language switch intent is detected, output ONLY "None".
- Valid outputs: , , ......... "None" (or other languages as specifically mentioned)
## Analysis Process:
* Language Switch Intent Detection:
- Check each engine output for explicit requests to change languages such as:
- English phrases: "switch to [language]", "change language to [language]", "I want to speak in [language]", "let's talk in [language]"
- Hindi phrases: "भाषा बदलें [language]", "मैं [language] में बात करना चाहता हूं", "चलो [language] में बात करें"
.
.
.
.
-
- Look for language-specific switch indicators:
- Words like "switch", "change", "speak in", "talk in" followed by a language name
- Similar phrases in other languages that indicate a desire to change languages
- For partial or unclear requests, cross-reference all STT outputs to determine the intent
* Target Language Identification:
- If a switch intent is detected, identify the target language requested:
- For direct mentions: Extract the language name from the phrase
- For implicit mentions: Infer from context which language is being requested
b. Only consider these valid target languages (unless others are specifically mentioned):
-
-
.
.
-
- Decision Rules:
a. If ANY of the engine outputs contain a clear language switch request:
Output the requested language name
b. If NONE of the engine outputs contain a language switch request:
Output "None"
c. If multiple conflicting language switches are requested:
Output the language that appears most consistently across engines
Sample Input/Output Examples:
Example 1:
Input:
S1: "ಮುಂದೆ ಕನ್ನಡದಲ್ಲಿ ಮಾತನಾಡೋಣ"
S2: "let's speak in Kannada"
S3: "ನಾವು ಕನ್ನಡದಲ್ಲಿ ಮಾತನಾಡೋಣ"
Output:
Kannada
Example 2:
Input:
S1: "తెలుగులో మాట్లాడుదాం"
S2: "let's talk in Telugu"
S3: "నేను తెలుగులో మాట్లాడాలనుకుంటున్నాను"
Output:
Telugu
IMPORTANT: The output must ONLY contain the target language name or "None". Do not include any analysis, explanation, or other text in the output.
```
### **Trigger Workflow**
When a new language is added to the agent via the UI:
1. A **confirmation dialog** is shown instructing you to update the Language Switch Prompt.
2. Navigate to Agent Overview → Language Switch Prompt.
3. Modify the reference prompt to include:
1. Any new language added
2. Detection examples/phrases (if required)
4. Save and deploy.
Failure to update this prompt will result in language switching not working, even if the language has been configured elsewhere.
### **Best Practices**
| **Practice** | **Recommendation** |
| :----------------- | :-------------------------------------------------------------------------------------------------- |
| Prompt Matching | Always ensure the languages listed in the prompt match the ones added to the agent’s system prompt. |
| Script Awareness | Include detection patterns in local scripts (e.g., Hindi, Kannada) if relevant. |
| Clear Output | Always return a clean one-word output. Avoid additional logs, quotes, or structures. |
| Prompt Maintenance | Revisit and update the prompt whenever you add, remove, or rename any supported languages. |
| Sample Inputs | Use actual transcriptions from user speech; avoid manually translated phrases for accuracy. |
# Using Jinja for Dynamic System Prompts
Source: https://docs.gnani.ai/C03_Jinja
Description of your new file.
### **Overview**
The Jinja templating feature allows you to create powerful, dynamic System Prompts. By using variables that are resolved with real-time data before a call begins, you can build personalized and context-aware conversations. This guide explains how to use Jinja templates, the validation rules in place, and the benefits of this approach.
## **Core Concept: Understanding Variable Types**
Our system uses two distinct types of curly braces to handle variables from different sources. Understanding this distinction is crucial for building prompts correctly.
### \*\*1. Server-Side Variables: \*\*
These variables are placeholders for data that you provide to the system *before* the call is initiated.
* **Source:** The value for these variables must come from one of two places:
* **Dynamic API Response:** Data fetched from your API endpoint just before the call.
* **Pre-call Variables:** Data you provide when triggering the call via an API, especially during testing.
* **Use Case:** Ideal for injecting customer-specific data like names, appointment details, order history, or account status.
* **Example:** Hello customer\_name , your appointment is scheduled for booking\_date .
### \*\*2. User-Input Variables: \*\*
These variables are placeholders for information that is extracted directly from the **user's spoken message** during the conversation.
* **Source:** The value is populated by the system based on the user's utterance.
* **Use Case:** Perfect for capturing dynamic user choices or inputs within a conversation turn.
* **Example:** If the prompt is “To confirm, you selected user\_selected\_option, right?”, the value of user\_selected\_option will be filled from what the user said in their previous turn.
### ***Variable Reference Table***
| Syntax | Example | Source of Value | Required in API/Pre-call Variables? |
| :------------- | :--------------------- | :----------------------------------------- | :---------------------------------- |
| variable\_name | customer\_name | Dynamic API response or Pre-call variable | **Yes** |
| variable\_name | user\_selected\_option | Extracted directly from the User's Message | **No** |
### **Validation Rules and System Behavior**
To ensure your prompts are reliable, the system validates the Jinja template syntax before you can save or test an agent.
***Validation Checks:*** The System Prompt is validated for the following conditions:
* **Correct Jinja Syntax:** Ensures all brackets and statements (e.g., `{{var}}`, `{ %if% }`) are correctly formatted.
* **No Undeclared Variables:** All variables wrapped in double curly braces `{{...}}` must have a corresponding value provided either from the dynamic API or pre-call variables.
### ***Error Handling***
The system's behavior changes based on the validation result:
| Condition | Save / Test Buttons | System Message |
| :------------------------------------------------------------------ | :--------------------- | :----------------------------------------------- |
| **Valid Syntax** | ✅ Enabled | – |
| **Invalid Jinja Syntax** | 🛑 Disabled | "Invalid Jinja syntax." |
| **Variable Missing at Runtime** (Variable not in API/pre-call data) | ✅ Enabled (Save works) | "Call triggered failed." (Error during the call) |
### ***Agent Behavior***
* **Existing Agents:** Agents created before this feature will continue to function as normal until their System Prompt is modified and saved again.
* **New & Updated Agents:** All new agents, or existing agents whose prompts are updated, must follow the Jinja formatting and validation rules.
### **Examples**
**1. Valid Syntax**
```html theme={null}
Hello {{customer_name}}, your booking for a {{product_name}} is confirmed for {{booking_date}}.
```
**2. Invalid Syntax**
```html theme={null}
Hello {{customer_name}}. Your booking is on {{booking_date}.
```
* **Error:** The curly braces for booking\_date are not closed.
* **System Behavior:** The Save and Test buttons will be disabled until the syntax is corrected.
**3. Advanced Example with Conditional Logic**
```html theme={null}
Hello {{ customer_name }}.
{% if loyalty_status == 'Gold' %}
As a Gold member, you get a special 20% discount.
{% elif loyalty_status == 'Silver' %}
As a Silver member, you get a 10% discount.
{% else %}
Thank you for being our customer.
{% endif %}
Your order #{{ order_id }} is ready.
```
*In this example, the agent's greeting is personalized based on the customer's loyalty\_status variable.*
### **How Jinja Optimizes Cost and Performance**
Using Jinja templates, especially with conditional blocks (`{% if %}`, `{% elif %}`, `{% else %}`), significantly optimizes your prompts.
* **Reduces Input Tokens:** By including only the relevant text based on the provided variables, the overall length of the prompt sent to the language model is reduced.
* **Lowers Inference Costs:** Fewer input tokens directly translate to lower API costs for each call.
* **Improves Execution Speed:** Shorter, more concise prompts are processed faster by the language model, reducing latency.
**Important Note:** During **testing mode**, if a variable exists in both the pre-call variables and the dynamic API response, the value from the **pre-call variables will be used**.
# SMS Integration and Action
Source: https://docs.gnani.ai/D01_Twilio_SMS
Enable your agent to send SMS notifications like appointment reminders, ticket confirmations, order updates, discount codes, and more — all directly from Gnani Agents using Twilio.
***
### **Why Integrate SMS?**
With SMS integration, your agent can:
* Send **appointment reminders** to reduce no-shows.
* Deliver **order updates** and **ticket confirmations** instantly.
* Share **promotional offers** or **discount codes** with customers.
***
## 1. Steps to Integrate
1. **Open Integration Settings**
* Go to \*\*Integration \*\*section in your Manage Agent page.
* Add a new integration. Select SMS.
* Select **Twilio** as the operator and enter a name for your integration.
2. **Get Twilio Credentials**
* Open the [Twilio Console Dashboard](https://console.twilio.com/dashboard) to retrieve credentials.
* Copy your **Account SID** (34- character alphanumeric identifier starts with `AC`) and **Auth Token** from the **Account Info** section.
* Paste these into the integration form in Gnani Agents.
3. **Choose Sender Type**
\*\*Option A — Phone Number \*\*to send messages from a fixed number for consistent branding.
* In Twilio: Go to **Develop → Phone Numbers → Manage → Active Numbers**.
* Select an active **SMS-enabled** number and paste it into Gnani Agents.
* If you don’t have one, click **Buy a Number** in Twilio.
\*\*Option B — Messaging Service \*\*to use a pool of numbers for sending SMS dynamically based on factors like location and compliance.
* In Twilio: Go to **Develop → Messaging → Services**.
* Copy the **Messaging Service SID** (starts with `MG`) and paste it into Gnani Agents.
* If you don’t have one, click **Create Messaging Service** in Twilio.
4. **Save**
* Click **Save** to complete the integration.
***
## **2. Create an SMS Action**
1. **Navigate to Actions**
* Go to **Manage Agents → Select Agent → Action Tab**.
2. **Create Action**
* Click **+** to open the **Create Action** card.
3. **Fill in Details**
* **Name:** Use only letters, numbers, or underscores.
* **Description:** Describe when the action should trigger.
* Example: *"Trigger this action when the user asks about available offers."*
4. **Select Integration**
* Choose the Twilio SMS integration you set up earlier.
5. **Choose Trigger**
* **Post-Call:** Executes after the conversation ends.
* If **Post-Call** is selected, you can:
* Add **Variables** (sent to the API).
* **On-Call:** Executes during the conversation.
* If **On-Call** is selected, you can:
* Write a message in the **Speak During Action** section (what the agent says when executing).
* Add **Before API Call Variables** (sent to the API). Create new variables or reuse existing ones to store dynamic values like names, order IDs, or dates.
* Add **After API Call Variables** (received from the API).
6. **Message Template**
* Enter the SMS content to send. Use variables for dynamic content.
* Example: `Hello {{user_name}}, your order {{order_id}} is on the way! `
***
## **3. Test Your SMS Action**
Before using the action in real conversations, test it to ensure it’s working correctly.
1. **Open Testing Mode**
* Go to **Manage Agents → Select Agent → Test**.
2. **Start a Test Conversation**
* Speak or chat with your agent inside the test interface.
3. **Trigger the Action**
* Ask the agent for similar instructions or keywords mentioned in the Action’s **Description**.
* Example: If your description says *"Trigger this when user asks about offers"*, say **"What offers do you have?"**.
4. **Check SMS Delivery**
* Confirm the SMS arrives at the configured number.
* Ensure placeholders (e.g., `{{user_name}}`) are replaced with actual values.
5. **Troubleshoot if Needed**
* Click on the more options on your created action card to see the **View Logs** option. You can check the **Action Logs** here.
* Verify Twilio credentials and sender configuration.
* Make sure the test number is SMS-enabled.
* Review triggers, message template, and variables for errors.
***
✅ **You’re ready!** Your agent can now send automated, personalized SMS messages during or after calls.
# CRM Integration and Action
Source: https://docs.gnani.ai/D02_Zoho_CRM
Integrate Zoho CRM with Gnani Agents to automatically create and manage customer records during or after conversations — making lead capture and support tracking seamless.
***
## **1. Why Integrate Zoho CRM?**
By connecting Zoho CRM to Gnani Agents, your agent can:
* **Automatically create leads, contacts, or deals** while talking to customers.
* Reduce manual data entry by **capturing customer details in real time**.
* Improve lead follow-up speed and support resolution times.
***
## **2. Set Up Zoho CRM Integration**
1. **Open Integration Settings**
* Go to **Add CRM Integration** in your Gnani Agents dashboard.
* Select **Zoho** as the operator and enter a name for your integration.
2. **Get Zoho Credentials**
* Open the Zoho API Console.
* Create a **Self Client** by following Zoho’s Self Client guide.
3. **Copy Client ID & Client Secret**
* In the **Client Secret** tab, locate your **Client ID** and **Client Secret**.
* Paste both into the integration form in Gnani Agents.
4. **Enter Account Server URL**
* Choose the correct **Account Server URL** based on your primary region from Zoho’s region list.
5. **Generate Account SOID**
* Follow **Step 1** in this Zoho guide to generate your **Account SOID**, then paste it into Gnani Agents.
6. **Specify Required Scopes**
* From Zoho’s scope list, add the required scopes for your action (e.g., `ZohoCRM.modules.ALL`).
7. **Save**
* Click **Save** to complete the integration.
***
## **3. Create a CRM Action**
1. **Navigate to Actions**
* Go to **Manage Agents → Select Agent → Action Tab**.
2. **Create Action**
* Click **+** to open the **Create Action** card.
3. **Fill in Details**
* **Name:** Use only letters, numbers, or underscores.
* **Description:** Describe when the action should trigger.
* Example: *"Trigger this action when the user provides new lead details."*
4. **Select Integration**
* Choose your **Zoho CRM** integration.
5. **Choose Trigger**
* **Post-Call:** Executes after the conversation ends.
* If **Post-Call** is selected, you can:
* Add **Variables** (sent to the API).
* **On-Call:** Executes during the conversation.
* If **On-Call** is selected, you can:
* Write a message in the **Speak During Action** section.
* Add **Before API Call Variables** (sent to the API).
* Add **After API Call Variables** (received from the API).
6. **Set CRM-Specific Fields**
* **Select Action:** Choose the CRM action type (e.g., `CREATE` for adding new records).
* **Module Name:** Specify the CRM module where the record should be created (e.g., `Leads`, `Accounts`, `Contacts`, `Deals`).
* **Payload:** Define the data in JSON format. Example:
```json theme={null}
{
"Company":"Example Corp",
"Name": "Smith",
"Email": "john.smith@example.com"
}
```
***
## **4. Test Your CRM Action**
Testing ensures the integration works correctly before using it in live conversations.
1. **Open Testing Mode**
* Go to **Manage Agents → Select Agent → Test**.
2. **Start a Test Conversation**
* Chat or speak with your agent in the test interface.
3. **Trigger the Action**
* Use the same instructions or keywords from your Action’s **Description** to activate it.
4. **Verify in Zoho CRM**
* Log in to Zoho CRM and check if the record was created or updated as expected.
5. **Check Action Logs**
* In Gnani Agents, click **More Options (⋮)** on the Action card → **View Logs**.
* Review execution details and any API errors.
***
✅ **Your Zoho CRM integration is now ready!** Your agent can capture leads and customer details without manual effort.
# Email Integration and Action
Source: https://docs.gnani.ai/D03_MailSend_Email
Allow agents to send emails
Integrate your email provider with Gnani Agents to automatically send personalized emails — such as welcome messages, follow-ups, newsletters, and transactional notifications — without manual effort.
***
## **1. Why Integrate Email?**
By connecting an email service to Gnani Agents, your agent can:
* **Send automated welcome emails** when a new customer signs up.
* Deliver **follow-up sequences** after calls.
* Share **newsletters, offers, or promotions** at the right moment.
* Trigger **transactional emails** (order confirmations, invoices, etc.) instantly.
***
## **2. Set Up Email Integration**
You can integrate with either **Mailchimp** or **SendGrid** depending on your preference.
***
### **Option 1 — Mailchimp Integration**
1. **Open Integration Settings**
* Go to **Add Email Integration** in your Gnani Agents dashboard.
* Select **Mailchimp** as the operator and provide an integration name.
2. **Get Your API Key**
* Log in to your Mailchimp account.
* Go to **Account → Extras → API Keys**.
* Create or copy an existing **API Key**.
3. **Add API Key to** Gnani Agents
* Paste the Mailchimp API Key into the integration form in Gnani Agents.
4. **Save**
* Click **Save** to complete the integration.
***
### **Option 2 — SendGrid Integration**
1. **Create or Log Into Your Account**
* Go to [SendGrid](https://sendgrid.com) and sign up or log in.
2. **Create an API Key**
* In your SendGrid dashboard, go to **Settings → API Keys**.
* Click **Create API Key**.
* Name it (e.g., *Platform Integration*).
* Set **API Key Permissions** to **Full Access** or at least **Mail Send**.
* Click **Create & View** — copy this key (you won’t see it again later).
* Paste it into Gnani Agents' **API Key** field.
3. **Add & Verify Your Sender Email**
**Option A — Single Sender Verification (Quick)**
* Go to **Settings → Sender Authentication → Verify Single Sender**.
* Enter your sender email (e.g., `support@yourdomain.com`) and details.
* Click the verification link sent to your email.
* Use this sender email in Gnani Agents.
**Option B — Domain Authentication (Recommended for Production)**
* If you manage your domain’s DNS, choose **Domain Authentication**.
* SendGrid will give you **CNAME records** to add to your DNS.
* Once verified, you can send from any address on that domain.
4. **Save**
* Click **Save** to complete the integration.
***
## **3. Create an Email Action**
1. **Navigate to Actions**
* Go to **Manage Agents → Select Agent → Action Tab**.
2. **Create Action**
* Click **+** to open the **Create Action** card.
3. **Fill in Details**
* **Name:** Use only letters, numbers, or underscores.
* **Description:** Describe when the action should trigger.
* Example: *"Trigger this action when the user provides their email for newsletter signup."*
4. **Select Integration**
* Choose your Mailchimp or SendGrid integration.
5. **Choose Trigger**
* **Post-Call:** Executes after the conversation ends.
* If **Post-Call** is selected, you can:
* Add **Variables** (sent to the API).
* **On-Call:** Executes during the conversation.
* If **On-Call** is selected, you can:
* Write a message in the **Speak During Action** section.
* Add **Before API Call Variables** (sent to the API).
* Add **After API Call Variables** (received from the API).
6. **Set Email-Specific Fields**
* **Select Action:** Select the default option *(Contact us for support for multiple Mailchimp actions)*
* **Template Name:** Select a Mailchimp email template to use.
* **From Name & Email:** The verified sender name and email in Mailchimp.
* **To Name & Email:** The recipient’s details. Variable input is recommended (e.g., `{{customer_name}}`, `{{customer_email}}`).
* **Add Subject:** Custom subject line — can include variables.
* **Merge Variables:** Map Gnani Agents variables to Mailchimp template placeholders:
* **Key:** Variable name in Mailchimp template.
* **Value:** Corresponding Gnani Agents variable (e.g., `{{order_id}}`).
***
## **4. Test Your Email Action**
Testing ensures your emails are triggered correctly before going live.
1. **Open Testing Mode**
* Go to **Manage Agents → Select Agent → Test**.
2. **Start a Test Conversation**
* Speak or chat with your agent in the test interface.
3. **Trigger the Action**
* Use the same instructions or keywords from your Action’s **Description** to activate it.
4. **Check Your Email**
* Verify that the email is received by the intended recipient.
* Confirm that template variables are replaced correctly with actual data.
5. **Check Action Logs**
* In Gnani Agents, click **More Options (⋮)** on the Action card → **View Logs**.
* Review the logs for execution details and any API errors.
***
✅ **Your Email integration is now ready!** Your agent can send beautifully designed, personalized emails using Mailchimp or SendGrid directly from conversations.
# Ticketing Integration and Actions
Source: https://docs.gnani.ai/D04_Zoho_Ticket
Integrate Zoho Desk with Gnani Agents to automatically create and manage support tickets — helping your team respond faster and resolve customer issues efficiently.
***
## **1. Why Integrate Ticketing?**
By connecting Zoho Desk to Gnani Agents, your agent can:
* **Automatically log customer issues** during or after a call.
* Ensure no support requests are missed.
* Reduce manual ticket creation time for your support team.
***
## **2. Set Up Zoho Desk Integration**
1. **Open Integration Settings**
* Go to **Add Ticket Integration** in your Gnani Agents dashboard.
* Select **Zoho** as the operator and enter an integration name.
2. **Get Zoho Credentials**
* Open the Zoho API Console.
* Follow Zoho’s Self Client guide to create a Self Client.
3. **Copy Client ID & Client Secret**
* In the **Client Secret** tab, locate your **Client ID** and **Client Secret**.
* Paste both into the integration form in Gnani Agents.
4. **Enter Account Server URL**
* Choose the correct **Account Server URL** based on your primary region from Zoho’s region list.
5. **Generate Account SOID**
* Follow this Zoho guide to generate your **Account SOID** and paste it into Gnani Agents.
6. **Specify Required Scopes**
* From Zoho Desk’s scope list, add the scopes needed for your ticket actions (e.g., `ZohoDesk.tickets.CREATE`).
7. **Save**
* Click **Save** to finalize the integration.
***
## **3. Create a Ticket Action**
1. **Navigate to Actions**
* Go to **Manage Agents → Select Agent → Action Tab**.
2. **Create Action**
* Click **+** to open the **Create Action** card.
3. **Fill in Details**
* **Name:** Use only letters, numbers, or underscores.
* **Description:** Describe when the action should trigger.
* Example: *"Trigger this action when the user reports a technical issue."*
4. **Select Integration**
* Choose your Zoho Desk integration.
5. **Choose Trigger**
* **Post-Call:** Executes after the conversation ends.
* If **Post-Call** is selected, you can:
* Add **Variables** (sent to the API).
* **On-Call:** Executes during the conversation.
* If **On-Call** is selected, you can:
* Write a message in the **Speak During Action** section.
* Add **Before API Call Variables** (sent to the API).
* Add **After API Call Variables** (received from the API).
6. **Set Ticketing-Specific Fields**
* **Select Action:** Choose the ticketing action type (e.g., `CREATE` for new tickets). *(Contact us to add more actions)*
* **Department ID:** Specify the Zoho Desk department for ticket creation.
* **Contact ID:** Associate the ticket with a customer profile. *(Currently for testing; contact us to make it dynamic)*
* **Payload:** Define the ticket in JSON format. The **subject**, **description**, and **priority** are mandatory:
```json theme={null}
{
"subject": "{{subject}}",
"description": "{{description}}",
"priority": "{{priority}}"
}
```
***
## **4. Test Your Ticket Action**
1. **Open Testing Mode**
* Go to **Manage Agents → Select Agent → Test**.
2. **Start a Test Conversation**
* Speak or chat with your agent in the test interface.
3. **Trigger the Action**
* Use the same instructions or keywords from your Action’s **Description** to activate it.
4. **Verify in Zoho Desk**
* Check Zoho Desk to confirm that a ticket is created with the correct subject, description, and priority.
5. **Check Action Logs**
* In Gnani Agents, click **More Options (⋮)** on the Action card → **View Logs**.
* Review logs for execution details and any errors.
***
✅ **Your Zoho Desk integration is now ready!** Your agent can log support tickets instantly, ensuring fast and efficient customer service.
# Custom Integrations and Actions
Source: https://docs.gnani.ai/D05_Custom
Description of your new file.
Custom integrations allow you to connect **any external API** to your Gnani Agents — perfect for unique workflows and advanced automation that go beyond built-in integrations.
***
## **1. Why Use Custom Integrations?**
With a custom integration, you can:
* Connect your agent to any API endpoint you control or have access to.
* Extend agent capabilities to trigger workflows in other systems.
* Handle use cases not covered by predefined integrations.
***
## **2. Set Up a Custom Integration**
1. **Open Integration Settings**
* Go to **Add Custom Integration** in your Gnani Agents dashboard.
2. **Enter Basic Details**
* **Integration Name**: Use letters, numbers, or underscores only.
* **Description**: Briefly explain the integration’s purpose.
3. **Choose API Call Method & URL**
* Select the **Method** (e.g., `GET`, `POST`, `PUT`, `DELETE`).
* Enter the **API URL** where the request should be sent.
4. **Add Authentication (If Needed)**
* If your API requires authentication, enter the **Key** and **Value**.
* Example: `Key: Authorization Value: Bearer `
5. **Save Integration**
* Click **Integrate** to complete the setup.
***
## **3. Create a Custom API Action**
1. **Navigate to Actions**
* Go to **Manage Agents → Select Agent → Action Tab**.
2. **Create Action**
* Click **+** to open the **Create Action** card.
3. **Fill in Details**
* **Name:** Use only letters, numbers, or underscores.
* **Description:** Describe when the action should trigger.
* Example: *"Send lead details to the CRM when the user shares contact info."*
4. **Select Integration**
* Choose your custom integration from the dropdown.
* **Method & URL** will be pre-filled from the integration setup.
5. **Enable Request Components**
* Based on your API method, you can **turn on/off**:
* **Headers**: For sending additional information about your API request, such as authentication or metadata.
* Example:`Key: Authorization Value: Bearer `
* **Params**: For adding filters or modifying the request. You can add multiple key-value pairs.
* Example:`Key: age Value: 30 `will filter by age
* **Body**: Data to send in the request (usually JSON).
* Example:
```yaml theme={null}
{
"user_text": "{{user_text}}",
"sender_id": "{{sender_id}}"
}
```
*
6. **Set Timout**
* Define the maximum wait time (in seconds) before the request times out.
* If the server doesn’t respond in time, the request fails with a timeout error.
7. **Choose Trigger**
* **Post-Call:** Executes after the conversation ends.
* If **Post-Call** is selected, you can:
* Add **Variables** (sent to the API).
* **On-Call:** Executes during the conversation.
* If **On-Call** is selected, you can:
* Write a message in the **Speak During Action** section.
* Add **Before API Call Variables** (sent to the API).
* Add **After API Call Variables** (received from the API).
***
## **4. Test Your Custom API Action**
1. **Open Testing Mode**
* Go to **Manage Agents → Select Agent → Test**.
2. **Start a Test Conversation**
* Speak or chat with your agent.
3. **Trigger the Action**
* Use the same instructions or keywords from your Action’s **Description** to activate it.
4. **Check API Logs**
* In Gnani Agents, click **More Options (⋮)** on the Action card → **View Logs**.
* Review request/response details and confirm the API executed correctly.
5. **Validate in External System**
* If your API triggers a process (e.g., record creation), check the connected system to confirm results.
***
✅ **Your Custom API integration is now ready!** You can now expand your agent’s abilities to virtually any API-enabled service.
# Integrating Gnani Agents with Webex Contact Center
Source: https://docs.gnani.ai/D06_Webex
Leverage the power of Gnani Agents seamlessly within your Cisco Contact Center workflows. Gnani Agents integrates with Webex Contact Center through Cisco’s **Service App** framework. Once set up, you can use Gnani Agents inside the Webex **Flow Designer** to handle customer interactions seamlessly.
As an Admin, here’s what you need to do:
## Service App Creation: Admin Flow
### Step 1: Go to Service Apps in Control Hub
* Sign in to your Webex **Control Hub**.
* Navigate to **Management → Apps → Service Apps**.
### Step 2: Locate the Gnani Agents Service App
* You will see **Inya AI Agent** listed as a Service App available for your organization.
* Click on it to review the details and scopes requested.
### Step 3: Authorize the Service App
* Click **Authorize** to enable the Inya Service App.
* This grants the app the required permissions to function within your Webex environment.
### Step 4: Confirmation
* Once authorized, the Gnani team will complete the backend setup.
* You can now add Gnani Agents directly in the **Cisco Contact Center** **Flow Designer** to manage customer calls and chats.
## After Authorization
* No further configuration is needed from your side.
* The Gnani **team** will handle credit setup and ensure that your AI agents are active.
If you face any issues or need more credits, contact us at [**hello-inya@gnani.site**](mailto:hello-inya@gnani.site).
That’s it, with just a one-time authorization, your organization can start using Gnani Agents inside Webex Contact Center workflows.\
\
We'll assist you with:
* Purchasing and allocating required credits
* Completing backend setup and Webex deployment
* Verifying that your Gnani Agents are active within the Contact Center flow designer
## To add the Gnani Agents Transcript widget to the Desktop Layout:
1. Login to Cisco Webex Contact Hub Desktop as an Administrator
2. Click on **Contact Center** under **Services** on left pane, and then go to **Desktop Layout** under **Desktop Experience**
3. From the desktop Layouts page that opens, select the layout you want to edit.
4. Click on the download button next to the already uploaded JSON file to download so that you can edit it.
5. In the .json file, add the following details to create another tab for the Inya Transcript Widget to show up under “agent” > “area” > “panel” > “children” array
```json theme={null}
{
"comp": "md-tab",
"attributes": {
"slot": "tab",
"class": "widget-pane-tab"
},
"children": [
{
"comp": "slot",
"textContent": "Inya Transcript",
"attributes": {
"name": "INYA_TRANSCRIPT_TAB"
}
}
]
},
{
"comp": "md-tab-panel",
"attributes": {
"slot": "panel",
"class": "widget-pane"
},
"children": [
{
"comp": "dynamic-area",
"attributes": {
"name": "INYA_TRANSCRIPT"
},
"properties": {
"area": {
"id": "inya-transcript",
"widgets": {
"inya": {
"comp": "inya-transcript",
"script": "https://genvoice-appdev.gnani.site/sdk/inya/inya-webex-widget.min.js",
"properties": {
"agentId": "$STORE.agent.agentId",
"taskMap": "$STORE.agentContact.taskMap",
"agentName": "$STORE.agent.agentName",
"darkMode": "$STORE.app.darkMode",
"accessToken": "$STORE.auth.accessToken"
},
"wrapper": {
"title": "Inya Transcript",
"maximizeAreaName": "app-maximize-area"
}
}
},
"layout": {
"areas": [["inya"]],
"size": {
"cols": [1],
"rows": [1]
}
}
}
}
}
]
}
```
8. Once changes are done, save the file and upload the file in the Layout
9. Click **Save**
### Summary
By integrating Gnani Agents **with Webex**, you get all the advantages of intelligent, conversational AI embedded straight into your Cisco Contact Center. Just:
1. Set up via the Webex App Hub
2. Create your agents within Gnani Agents
3. Contact us at [**hello-inya@gnani.site**](mailto:hello-inya@gnani.site) for credit setup and activation
4. Insert your agents into Webex’s flow designer to power customer conversations
# Conversational Logs
Source: https://docs.gnani.ai/E01_Conversational_Logs
Your agent's flight recorder
Want to replay a user conversation or debug a tricky call? **Conversational Logs** are your time machine!
### What Are Conversational Logs?
Conversational Logs capture every conversation your agent has, including both test sessions and live interactions. They offer a detailed view of what happened during each call or chat, allowing you to review, analyze, and improve your agent’s performance.
### Why Are They Important?
* **Insightful Analysis:** Understand the flow of each conversation.
* **Performance Tracking:** Monitor call durations, latency, and overall performance.
* **Troubleshooting:** Identify issues like delays or miscommunications.
* **Quality Assurance:** Review call insights to assess resolutions and agent behavior.
### How to Access Conversational Logs
You have several ways to view these logs:
1. **For All Agents:**
* Go to **Agent Related → Conversational Logs** to see logs for all your agents together.
2. **For All Agent Chains:**
* Navigate to **Agent Chains → Conversational Logs** for a consolidated view of your agent chains.
3. **For a Specific Agent:**
* Go to **Agent Related → Manage Agents**, select an agent, then click **View Logs** at the top right.
4. **Post-Test Shortcut:**
* After ending a test conversation, a snack bar appears with a link to view that specific conversation log.
### What You’ll See in a Conversation Log
* **Start and End Times:** Know exactly when the conversation began and ended.
* **Latency:** Track any delays incurred during the conversation.
* **Call Insights:** Get summarized details on the call’s reason, result, overview, and resolution.
* **Recording and Transcript:** Listen to or read the complete conversation for in‑depth analysis.
# Dev Logs
Source: https://docs.gnani.ai/E02_Dev_Logs
Troubleshoot agent behavior
### **What Are Dev Logs?**
Dev Logs are a timestamped sequence of system-level events that occur during the lifecycle of a voice/chat interaction. These logs capture each stage of the agent orchestration, starting from the initial greeting to transcription, LLM processing, voice synthesis, and call disconnection. It's built for developers and advanced users who want visibility into every stage of the agent's thinking and speaking process.
### **Where to Find It**
You can access Dev Logs by navigating to:
**Conversation Logs → Dev Logs**
### **Why Use Dev Logs?**
Dev Logs are essential when:
* You want to debug why an agent didn't respond correctly
* You need to analyze latency at each stage (TTS, ASR, LLM)
* You're optimizing call performance
### **Log Source / Module Tags**
Each log line is prefixed by a module to indicate where it originated:
| **Module** | **Description** |
| :--------------- | :---------------------------------------------------------------------------------------------------------------------- |
| CALL | Tracks call lifecycle: start, disconnect |
| ORCH | Orchestration logic like greeting fetch, post-call actions, post-call processing |
| TTS | Text-to-speech processing, TTFB (time to first byte), latencies, audio durations |
| ASR / TRANSCRIBE | Speech recognition, transcription results, audio duration and latency |
| LLM | Large Language Model events, streaming, token usage info, language switch detection, error handling if generation fails |
| KB | Knowledge base context retrieval and metrics |
| BARGE | User interruption (barge-in) detection and handling |
| ERROR | Error responses, especially from OpenAI or other APIs |
| UNKNOWN | Fallback logs for unexpected events or raw debug info |
### **Latency Monitoring**
Look out for high latency values in:
* **ASR Latency**
* **TTS Latency**
* **LLM Generation Time**
High values may indicate network, model, or transcription delays. These are key areas to optimize.
### **End of Conversation Signals**
Look out for:
* EOC: End of Conversation
* CALL: Disconnecting call: User hang-up or timeout
* ORCH: Finalizing total call credits: Post-call processing
### **Debugging Tips**
| **Issue** | **Check This Section** |
| :---------------------------------------- | :--------------------- |
| Agent responded slowly | TTS / LLM latency |
| Agent misunderstood user | ASR final result |
| Agent gave a wrong or incomplete response | LLM logs |
| Response wasn’t spoken | TTS |
| Unexpected disconnect | CALL logs at the end |
### **Summary**
The Dev Logs tab is your deep-dive tool for full transparency into:
* What your agent said
* What the user said
* How long each step took
* What the model generated and why
It’s invaluable for debugging, improving accuracy, and ensuring top performance.
# Agent Analytics
Source: https://docs.gnani.ai/E03_Agent_Analytics
Your Agent’s Report Card
Is your agent chatty, efficient, or needs improvement? Find out here!
### What Are Agent Analytics?
Agent Analytics provides an overview of your agent’s performance through key metrics and trends. It’s your go-to tool for tracking how your agent interacts with users over time.
### Key Metrics You Can Monitor
* **Total Calls/Chats:** The overall number of interactions.
* **Total Duration:** The combined duration of all calls and chats.
* **Average Duration:** The average length of a conversation.
* **Connected Calls:** Calls successfully connected with users.
* **Sentiment Trend:** Insights into the emotional tone (positive, negative, neutral) of interactions over time.
* **Call Insights:**
* Top reasons: Reason of the call (e.g., refund requests)
* Intents: User intents identified in conversations.
* Topics: Frequently mentioned keywords (e.g., shipping delays)
* Dispositions: Call outcomes (e.g., resolved, unresolved, follow up requested)
* **Total Actions Triggered:** The number of integration actions executed.
* **Call Drop-off Rate:** The percentage of calls where users disconnected early.
### How to Access Agent Analytics
1. Navigate to **Agent Related → Manage Agents**.
2. Select the agent you want to analyze.
3. Click on **Analytics** at the top right.
### Why This Is Valuable
These analytics help you understand not just what your agent is doing, but how well it’s performing. Use these insights to fine-tune prompts, adjust integrations, and enhance the overall user experience.
# Action Logs
Source: https://docs.gnani.ai/E04_Action_Logs
Your Integrations’ Detective tool.
*"Did My SMS Send? Why Did the CRM Update Fail?" Action Logs are where you’ll find your answers.*
### What Are Action Logs?
Action Logs help you verify that the actions triggering integrations with external services (CRM, email, SMS, etc.), are working correctly. They provide detailed records of every action executed during conversations.
### Why Use Action Logs?
* **Functionality Check:** Confirm that your integrations are performing as expected.
* **Detailed Feedback:** View trigger time, payload data, response, and status codes.
* **Troubleshooting:** Quickly pinpoint and resolve any issues with action execution.
### How to Access Action Logs
1. Navigate to **Agent Related → Manage Agents**.
2. Select the specific agent you want to review.
3. Go to the **Action** tab.
4. Click the kebab menu (three vertical dots) for the action you’re interested in.
5. Select **View Logs** to see the triggered actions list.
### What to Look For
* **Triggered Time:** When the action was initiated.
* **Payload:** The data sent to the external service.
* **Response:** Feedback from the service.
* **Status Code:** Code indicating success or type of error.
* **Parameters:** Additional details or variables used in the action.
# Need Help?
Source: https://docs.gnani.ai/F01_Support
Get Help, Share Feedback, and Report Issues
The **Support Desk** allows you to quickly raise tickets for bug reports, feature requests, account issues, and general feedback. Whether you need technical assistance or want to suggest an improvement, this section ensures your concerns are addressed efficiently.
***
## How to Raise a Support Ticket
### Step 1: Open the Support Desk
* Navigate to the **Support Desk** option in the sidebar.
* Click to open the support ticket submission form.
### Step 2: Fill in Ticket Details
The form contains five key fields to provide relevant information:
1. **Raising Ticket As:**
* Your registered email is auto-filled in this section.
* This field is non-editable, ensuring all tickets are linked to the logged-in user.
2. **Ticket Title:**
* Enter a short, clear title describing the issue or request.
3. **Category:**
* Choose from the following categories:
* **Bug Report:** Reporting errors, glitches, or unexpected behavior.
* **Feature Request:** Suggesting a new feature or improvement.
* **Account Issue:** Problems related to login, billing, or account settings.
* **Feedback:** General comments or suggestions.
* **Others:** Anything that doesn’t fit the above categories.
4. **Description:**
* Provide a detailed explanation of the issue or request.
* Include steps to reproduce the issue (if applicable) and any relevant details.
5. **Upload Screenshots (Optional):**
* Click to upload or drag and drop up to **three** image files.
* Supported formats: **JPG, JPEG, PNG** (Max size: **5MB per file**).
### Step 3: Submit the Ticket
* Once all details are filled in, click **Submit**.
* Your ticket will be logged, and the support team will review it promptly.
***
## Best Practices for Raising a Ticket
✔ **Be Specific:** Provide clear and concise information.
✔ **Attach Screenshots:** Helps the support team understand the issue faster.
✔ **Use Relevant Categories:** Ensures your ticket is handled efficiently.
✔ **Check for Help Docs:** Check for solutions in the help docs section.
Once submitted, you can expect timely updates on your ticket's progress. Need urgent help? Look out for additional support channels within the platform!
🚀 **Your feedback helps us improve, don’t hesitate to reach out!**
# Organizations
Source: https://docs.gnani.ai/F02_Org
This guide explains how Organizations, Roles, Agent Access, Sharing, and Environments work in Gnani Agents.
It is designed for teams that want structured collaboration, controlled access, and a safe promotion workflow before going live.
***
## Who Should Use Organizations?
Organizations are ideal for:
* Teams building agents collaboratively
* Companies separating development and testing responsibilities
* Businesses that need controlled access to agents
* Teams that want structured Development → Staging → Production workflows
If you are working alone, you may use your Personal workspace. If you are working with a team, using an Organization is strongly recommended.
***
## Why Use an Organization?
Using an Organization provides:
* Clear role-based access (Developer, QA, Org Admin)
* Controlled agent visibility
* Structured testing before production readiness
* Centralized ownership and oversight
* Safe collaboration without accidental overwrites
Organizations ensure that only the right people can build, test, and prepare agents for deployment.
***
## 1. Understanding Organizations
### What is an Organization?
An Organization is a shared workspace where:
* Team members collaborate
* Agents are created and managed
* Agents move through Development → Staging → Production
Deployment to live infrastructure is handled by the Gnani Agents team. Agents must be in the Production environment before deployment can be requested.
### How to Create an Organization
Organizations are not created through the UI.
If you need an Organization set up, please contact us and we will create it for you. We will also assign the initial Org Admin from our end.
If you need additional Org Admins added later, you can contact us for assistance.
***
## 2. Roles
An Organization includes three user roles:
* Org Admin
* Developer
* QA
Permissions are currently fixed. Custom roles and granular permission editing are coming soon.
| Role | Primary Responsibility |
| :-------- | :----------------------------------------- |
| Org Admin | Manage team members and oversee all agents |
| Developer | Build and manage agents in Development |
| QA | Test agents in Staging and mark them ready |
***
## 3. Role Responsibilities
### Org Admin
The Org Admin oversees the organization.
Can:
* Add/Remove Developers and QA users
* Change user roles (Developer ↔ QA)
* View and edit all agents
* Promote agents through environments
### Developer
Developers are responsible for building agents.
Can:
* Create agents in Development
* Edit agents in Development
* Share agents with other Developers
* Promote agents to Staging
Cannot:
* Edit agents in Staging
* Edit agents in Production
* Deploy agents
### QA
QA users validate agents before production readiness.
Can:
* View agents in Staging
* Test agents
* Mark agents as "Ready for Production"
Cannot:
* Edit agents
* Deploy agents
* Access Development unless granted visibility through the environment workflow
***
## 4. Membership Concept (Agent Access)
Agent access works on a membership model.
You can think of this as:
* Agent Access List
* Agent Membership
* Agent Ownership & Sharing
When a Developer creates an agent:
* That Developer automatically has membership
* The Org Admin automatically has membership
No one else can see or access that agent unless it is explicitly shared.
This ensures agents are private by default and only visible to intended collaborators.
### Sharing Agents
Agents can be shared with other Developers within the Organization.
When sharing an agent, there are two access levels:
* Read-Only Access
* Edit Access
### Edit Access Rules
* Only the **bot owner** (the Developer who created the agent) can grant **edit access** to another user.
* A user who is **not the bot owner** can only share the agent as **read-only**.
This ensures ownership control while still enabling collaboration.
Once shared:
* Read-only users can view the agent configuration but cannot modify it.
* Users with edit access can make changes in the Development environment.
QA users do not receive access through manual sharing.\
QA visibility is granted automatically when the agent is promoted to Staging.
### Deletion Rules for Shared Agents
If the agent creator (Developer A) shares an agent with another Developer (Developer B):
* If Developer B deletes the agent from their view, only their **membership** is removed.
* The agent itself is **not deleted** from the Organization.
* The agent remains visible to the bot owner, Org Admin, and any other users who have access.
Only the bot owner can permanently delete the agent from the Organization.
***
## 5. Agent Visibility Rules
### Default Visibility
When a Developer joins an Organization:
They see no agents by default.
They only see:
* Agents they created
* Agents shared with them
### Org Admin Visibility
Org Admins have visibility into all agents within the organization.
### QA Visibility
QA users:
* Do not see Development agents
* Automatically gain visibility when an agent is promoted to Staging
This ensures structured separation between building and testing.
***
## 6. Personal vs Organization Workspace
Every user has a Personal workspace.
Organization workspaces are shared team environments.
Currently:
* Transfer of agents from Personal → Organization is not available
* Agents cannot be moved between organizations
Agent transfer capabilities are coming soon.
Agents must be created directly within the intended workspace.
***
## 7. Environments & Workflow
The workflow follows a strict linear structure:
Development → Staging → Production
There is currently:
* No version history
* No rollback
* No skipping environments
* Overwrite model (each promotion replaces the previous configuration)
Versioning and advanced release management capabilities are coming soon.
### Development
Access:
* Developer
* Org Admin
Purpose:
* Build and edit the agent
* Modify prompts, flows, configurations
Action Available:
* Promote to Staging
Effect:
* Development configuration is copied to Staging
* Any existing Staging version is overwritten
### Staging
Access:
* QA (full testing)
* Org Admin (testing)
* Developer (read-only)
Purpose:
* Validate agent behavior
* Conduct testing
Action Available (QA):
* Mark "Ready for Production"
No editing is allowed in Staging.
### Production
Access:
* Org Admin (view and manage status)
* Other users (view only)
Production represents the finalized configuration.
Deployment is handled by the Gnani Agents team. An agent must be in Production before deployment can be requested.
***
## 8. Deployment Process
Deployment to live infrastructure is not triggered directly from the platform.
Once:
* The agent is in Production
* Testing is complete
You can contact the Gnani Agents team to initiate deployment.
Only agents in Production are eligible for deployment.
***
## 9. Integration Deletion Safety
If a Developer removes an integration from their agent:
* It is removed from their access
* It remains in the system if another agent is using it
An integration cannot be permanently deleted if it is linked to any active agent.
This prevents accidental disruption of other team members’ agents.
***
## 10. Current Limitations
The following capabilities are not yet available but are coming soon:
* Audit logs
* Version history and rollback
* Self-serve organization creation
***
## Summary
Organizations provide:
* Structured team collaboration
* Clear role separation
* Controlled agent visibility
* Safe Development → Staging → Production workflow
* Managed deployment process through the Gnani Agents team
This ensures predictable collaboration, clean testing flows, and controlled production readiness.
# Environments
Source: https://docs.gnani.ai/F03_Env
Environments allow teams to build, test, and prepare agents for deployment in a structured and controlled manner.
This feature is available only within Organizations and is designed to separate development, validation, and production readiness.
***
## Overview
Each agent inside an Organization has three environments:
* Development
* Staging
* Production
These environments follow a strict linear workflow:
Development → Staging → Production
Important:
* There is no version history.
* Saving changes overwrites the current state of that environment.
* Promoting to the next environment overwrites the configuration in that target environment.
Versioning and rollback capabilities are coming soon.
***
## 1. Development Environment
### Who Has Access?
* Developer
* Org Admin
### What Can Be Done Here?
Full editing capabilities are available:
* Modify system prompts
* Edit conversation flows
* Update knowledge base
* Change configurations
* Adjust integrations
This is the primary workspace where agents are built and iterated.
### Action: Promote to Staging
When you click **"Promote to Staging"**:
* The current Development configuration is copied to Staging.
* Any existing Staging configuration is overwritten.
* QA users automatically gain visibility of the agent in Staging.
This ensures QA only sees agents that are ready for testing.
***
## 2. Staging Environment
Staging is used for validation and testing before moving to Production.
### Who Has Access?
* QA
* Org Admin
### What Can Be Done Here?
QA can:
* Test conversations
* View logs
* Validate behavior
Editing is not allowed in Staging.
### Action: Mark "Ready for Production"
Available to: QA
When QA clicks **"Mark Ready for Production"**:
* The agent status changes to **Ready for Production**.
This signals that the agent has passed testing and is ready for final approval.
***
## 3. Production Environment
Production represents the finalized configuration of the agent.
### Who Has Access?
* Org Admin
### What Can Be Done Here?
* The agent handles live traffic once deployed.
* Editing is not allowed in Production.
### Action: Push to Production
Available to: Org Admin only
When **"Push to Production"** is clicked:
* The Staging configuration is copied to Production.
* Any existing Production configuration is overwritten.
This ensures that only a QA-approved configuration reaches Production.
***
## How Environment Promotion Works
Each promotion step fully replaces the target environment.
For example:
* Promoting from Development to Staging replaces everything in Staging.
* Pushing from Staging to Production replaces everything in Production.
Because there is no version history, previous configurations cannot be restored automatically.
***
## Deployment
Production represents the final approved configuration.
Deployment to live infrastructure is handled by the Gnani Agents team.
Only agents that are in the Production environment are eligible for deployment.
***
## Best Practices
* Complete all development work before promoting to Staging.
* Ensure QA thoroughly tests before marking Ready for Production.
* Confirm readiness before pushing to Production, since the action overwrites the existing configuration.
***
## Summary
The Environments feature ensures:
* Clear separation between building and testing
* Controlled approval before Production
* Reduced risk of accidental live changes
* Structured collaboration within Organizations
This workflow helps teams maintain quality, control, and clarity as agents move toward deployment.
# Setting Up Integrations
Source: https://docs.gnani.ai/M03_Integrations
Description of your new file.
## **Introduction**
Integrations allow your GenAI agent to connect with external systems, automating workflows and enhancing functionality. With integrations, your agent can **send emails, SMS messages, update CRMs, and manage tickets** effortlessly.
To use these integrations, you’ll need to configure **Actions**, which will be covered in the next module.
📍 **To add an integration:**
Navigate to **Manage Agents** → Select Agent → **Integration** tab → Add integration and select the required integration type.
***
## **Lesson 1: SMS Integrations**
### **Why Integrate SMS?**
Enable your agent to send appointment reminders, ticket confirmations, discount codes, and more via SMS.
### How to Set Up:
1. In the Add SMS Integration section, select Twilio as the operator, and provide a name for your integration.
2. Open the **Twilio Console Dashboard** ([https://console.twilio.com/dashboard](https://console.twilio.com/dashboard)) to retrieve credentials.
3. Copy the **Account SID** and **Auth Token** from the Account Info section and paste them into the platform.
* **Account SID:** A unique 34-character alphanumeric identifier starting with "AC".
* **Auth Token:** A secure key for authentication.
4. Choose the Sender type:
* **Phone Number:** Send messages from a fixed number for consistent branding.
* In Twilio, go to **Develop → Phone Numbers → Manage → Active Numbers**, select an active SMS-enabled number, and paste it into the platform.
* If you don’t have a number, click **Buy a Number** in Twilio.
* **Messaging Service:** Use a pool of numbers for sending SMS dynamically based on factors like location and compliance.
* In Twilio, navigate to **Develop → Messaging → Services**, find the **Messaging Service SID** (starting with "MG"), and paste it into the platform.
* If you don’t have a messaging service, click **Create Messaging Service** in Twilio.
5. Click **Save** to complete the integration.
***
## **Lesson 2: Email Integrations**
### **Why Integrate Email?**
Enable your agent to send **welcome emails, follow-ups, newsletters, and transactional messages** automatically.
### **How to Set Up:**
1. In the **Add Email Integration** section, select **Mailchimp** as the operator and provide an integration name.
2. Enter the **Mailchimp API Key** (found in your Mailchimp account).
3. Click **Save** to finalize the integration.
***
## **Lesson 3: CRM Integrations**
### **Why Integrate CRM?**
Allow your agent to **automatically create customer records** in Zoho CRM, streamlining lead and support management.
### **How to Set Up:**
1. In the **Add CRM Integration** section, select **Zoho** as the operator and enter an integration name.
2. Open the [Zoho API Console](https://api-console.zoho.in/) to retrieve credentials.
3. Follow [Zoho's guide](https://www.zoho.com/accounts/protocol/oauth/self-client/overview.html) to create a **Self Client**.
4. In the **Client Secret** tab, find the **Client ID** and **Client Secret** and paste them into the platform.
5. Enter the **Account Server URL** based on your primary user region ([choose from here](https://www.zoho.com/accounts/protocol/oauth/multi-dc.html)).
6. Generate the **Account SOID** using step 1 in [this guide](https://www.zoho.com/accounts/protocol/oauth/self-client/authorization-code-flow.html).
7. Specify the required scopes from [Zoho's scope list](https://www.zoho.com/crm/developer/docs/api/v3/scopes.html).
8. Click **Save** to complete the integration.
***
## **Lesson 4: Ticketing Integrations**
### **Why Integrate Ticketing?**
Enable your agent to **create support tickets** in Zoho Desk, automating customer service processes.
### **How to Set Up:**
1. In the **Add Ticket Integration** section, select **Zoho** as the operator and enter an integration name.
2. Open the [Zoho API Console](https://api-console.zoho.in/).
3. Follow [Zoho's guide](https://www.zoho.com/accounts/protocol/oauth/self-client/overview.html) to create a **Self Client**.
4. Retrieve the **Client ID** and **Client Secret** from the **Client Secret** tab and paste them into the platform.
5. Enter the **Account Server URL** ([choose from here](https://www.zoho.com/accounts/protocol/oauth/multi-dc.html)).
6. Generate the **Account SOID** using [this guide](https://www.zoho.com/accounts/protocol/oauth/self-client/authorization-code-flow.html).
7. Specify the required scopes from [Zoho Desk's scope list](https://desk.zoho.com/DeskAPIDocument#OauthTokens#OAuthScopes).
8. Click **Save** to finalize the integration.
***
## **Lesson 5: Custom Integrations**
### **Why Use Custom Integrations?**
For unique use cases, integrate **custom APIs** to expand your agent’s capabilities beyond built-in options.
### **How to Set Up:**
1. In the **Add Custom Integration** section, enter a name and description.
2. Choose an **API Call Method** (GET, POST, etc.) and enter the **API URL**.
3. If authentication is required, enter the **Key** and **Value**.
4. Click **Integrate** to complete the setup.
***
*🚀Your GenAI agent is now integrated and ready to automate workflows ! 🌍*
# Managing Variables & Actions
Source: https://docs.gnani.ai/M04_Actions
Description of your new file.
## **Introduction & Context**
Variables and actions power automation in your genAI agent. **Variables** allow dynamic data handling, while **actions** trigger workflows based on user interactions.
***
## **Lesson 1: Understanding Variables**
### **What Are Variables?**
Variables act as placeholders for storing and managing data throughout a conversation. They can:
* **Capture user input** (e.g., name, phone number, order details).
* **Store API responses** (e.g., retrieving order status). *(Coming Soon)*
* **Pass data into API requests** (e.g., including an ID for fetching details).
### **How to Create a Variable?**
1. **Define a Name** – Use letters, numbers, or underscores (e.g., `user_name`, `order_id`).
2. **Choose a Data Type**:
* **Boolean** (True/False)
* **Integer** (Whole numbers)
* **Float** (Decimal numbers)
* **String** (Text, such as names or messages)
3. **Set a Prompt** – Define how the variable gets its value.
* Example: "Ask the user for their PIN code (must be 6 digits)."
### **When to Use Variables?**
* **On-Call Triggering:**
* **Before an API Call** – Store user input to pass into an API request.
* **After an API Call** – Save data received from an API response. *(Clarification needed: What happens if the data isn’t stored?)*
* **Post-Call Triggering:** Use the variable to execute actions after a conversation.
### **Use Cases**
Personalizing emails, customizing SMS messages, or logging details in CRM systems.
***
## **Lesson 2: Setting Up Actions**
### **What Are Actions?**
Actions define **triggers** that send automated responses or execute integrations based on specific conditions.
### **How to Create an Action?**
1. Go to **Manage Agents → Select Agent → Action Tab**.
2. Click the **"+"** button to open the "**Create Action**" card.
3. **Enter a Name** - Use only letters, numbers, or underscores.
4. **Describe the Action** – Example: "Trigger this action when the user asks about available offers."
5. **Select an Integration** – Choose the service to execute the action.
6. **Choose a Trigger:**
* **Post-Call** – Executes after the conversation ends.
* **On-Call** – Executes during the conversation.
7. **Create or Add Variables (Optional):**
* Use variables to store dynamic data like customer names or order details.
* You can create new variables or reuse existing ones.
### **Integration-Specific Fields**
The fields below depend on the type of integration used.
### **For SMS Action:**
* **Message Template:** Define the SMS content. Insert variables using curly braces:
* Example: `"Hello {{user_name}}, your order {{order_id}} is on the way!"`
### For Email Action:
* **Select Action:** *(Work in Progress)*
* **Template Name:** Select an email template from Mailchimp.
* **From Name & Email:** The sender’s verified email address in Mailchimp.
* **To Name & Email:** The recipient's email (variable input is recommended).
* **Add Subject:** Customizable subject line (can include variables).
* **Merge Variables:** Map platform variables to Mailchimp template variables:
* **Key:** Variable in Mailchimp template.
* **Value:** Corresponding variable created in Inya (use curly braces).
***
### For CRM Action:
* **Select Action:** Choose a CRM action (e.g., CREATE for adding new records). *(Work in Progress)*
* **Module Name:** Specify the CRM module (e.g., Leads, Accounts, Contacts, Deals).
* **Payload:** Define JSON data format:\{ "Company": "Example Corp", "name": "Smith", "Email": "[john.smith@example.com](mailto:john.smith@example.com)"}
### For Ticket Action:
* **Select Action:** Choose a ticketing action (e.g., CREATE for new support tickets). *(Work in Progress)*
* **Department ID:** Specify the Zoho Desk department where the ticket should be created.
* **Contact ID:** Associate the ticket with a customer profile. Currently this is for testing the capabilities of the platform, but it will be a dynamic field in the future.
* **Payload:** Define the ticket structure in JSON format. The subject, description and priority are mandatory fields to create a ticket in Zoho Desk.
\{
"subject": "\{\{subject}}",
"description": "\{\{description}}",
"priority": "\{\{priority}}"
}
### For Custom API Action:
* **Method & URL:** Pre-populated based on the integration.
* Options to turn headers, params and body on and off will be available based on the API Method.
* **Headers:** Send additional information about your API request, such as authentication details or metadata:
* Example:
* **Key:** `Authorization`
* **Value:** `Bearer `
* **Params:** Params help filter or modify the request. You can add multiple key-value pairs.
* Example: Filter by age:
* **Key:** `age`
* **Value:** `30`
* **Body:** The body contains the data sent to the API, which is usually in JSON format. You can add multiple key-value pairs.
* **Key**: `name` → **Value**: `Jack Reacher`
* **Key**: `email` → **Value**: `reacher@example.com`
* **Timeout:** Define max wait time before the request times out. The platform will wait for a response from the server for the defined time before canceling the request. If the server doesn’t respond within this time, the request fails with a timeout error.
*Your genAI agent is now equipped with powerful automation tools! 🎯*
## **Lesson 2: Testing your Actions**
# Agent Chaining
Source: https://docs.gnani.ai/M05_Chains
Description of your new file.
The Agent Chaining feature on the Inya platform empowers you to build sophisticated, multi-step conversational workflows that truly understands and adapts to your users’ needs. With this, you can design a dynamic conversation pipeline where various nodes interact, decide, and make intelligent decisions. This guide will help you craft your own interactive flow!
## Lesson 1: What is Agent Chaining?
**Agent Chaining** is a visual and intuitive way to design conversation flows. Imagine it as a flowchart where each box (or node) represents a different part of the conversation. By connecting these nodes on a freeform canvas, you create a structured flow where each node handles a specific action or decision point.
### Why build an Agent Chain?
Imagine your genAI agent as a chef: a single recipe (prompt) makes one dish, but **Agent Chaining** lets you create a *full menu* with appetizers, mains, and desserts!
Similarly, Agent Chaining transforms a single specialized agent into a full-scale organization, with genAI ‘employees’ working across departments to solve complex business challenges.
A single-prompt conversational agent may struggle with complex, multi-step interactions. A large system prompt can cause it to lose track of events. With Agent Chaining, you overcome these limitations and gain:
* **Greater Control:** Guide your agent’s conversations through a predefined flow.
* **Contextual Decisions:** Decision nodes enable responses based on user input and conditions.
* **Enhanced User Experience:** Each node specializes in a purpose, making interactions smoother, natural and accurate.
* **Specialization:** Different knowledge bases or actions can be assigned to specific steps.
* **Avoid Pitfalls:** Prevent your agent from looping and hallucination by defining clear pathways.
**Use Cases:**
* **Customer Support:** The agent follows structured troubleshooting steps and escalates to a human agent when needed.
* **Sales:** Qualify leads with conditional questions and transfer calls when all criteria are met.
* **Business Owners:** Create customized conversation flows for various industries like healthcare, legal, and finance, ensuring every type of query is handled accurately.
***
## Lesson 2: Understanding the Node Types
### What Are Nodes?
Nodes are building blocks of your conversation flow. Each node represents an independent agent, by linking them together and defining conditions, you design the conversation path your agent chain will follow.
### Node Types Explained
1. **Default Node**
* **Purpose:** You can define the purpose of this node to your liking. It handles general queries and serves as a fallback when no other node fits.
* **How to Use:**
* Define a clear prompt for the agent’s response.
* Enable the “Static Text” if you want a fixed response.
* **Tip:** Use detailed context-rich prompts to ensure natural, effective responses.
2. **User Node**
* **Purpose:** Captures input from the user.
* **How to Use:** Simply add a User node to your flow where you need to collect information like a name, email, or specific responses.
3. **Decision Node**
* **Purpose:** Determines the next step in the conversation based on set conditions.
* **How to Use:**
* **Decision Prompt:** Add a condition to choose that path. (E.g., User’s age is less than 18)
* **Condition:** Optionally, define a prompt in the decision node for complex conditions.
* **Repeat Count:** Toggle and set a limit for allowing reattempts.
* **Knowledge Base:** Link a knowledge base for contextual decision-making.
4. **Event Node**
* **Purpose:** Triggers key events like ending the call, transferring the call to a human agent, or resetting the flow.
* **How to Use:** Select an event type (such as End Call, Reset, or Transfer) and define the response type (dynamic prompt or a fixed message).
**Connection Rules:**
* **Default** → Can connect to **User** nodes.
* **User** → Can connect to **Decision**, **Default**, or **Event** nodes.
* **Decision** → Can connect to **Default** or **Event** nodes.
* **Event** → Dead-end (conversation stops, restarts or gets transferred).
***
## Lesson 3: **Creating Your First Agent Chain**
Now that you understand the basics, it’s time to create your very first agent chain.
You can create an agent chain in three ways: from scratch, using a prompt, or uploading a JSON file.
### Lesson 3a: Creating from Scratch
1. **Start Fresh:**
* Go to ‘**Agent Chains**’ and click ‘**+**’.
* Select ‘**Create from Scratch**’, name your chain, choose an industry, and pick an icon.
2. **Begin with the Start node:**
* A Start and User node are automatically created.
* Customize the start node with a greeting message (static or prompt explaining how to start the conversation).
3. **Build Your Flow:**
* Add more nodes as required by clicking on the ‘**Add Node**’ option at the bottom.
* **Quick Add:** Hover over an existing node to reveal the connection points on the sides. Click on one to see a menu of compatible nodes that can be added and automatically connected to your selected node.
* For default and event nodes, choose if the node should speak a static text or function dynamically according to the prompt.
* For decision nodes, set up your conditions in the path and also the decision node.
* Connect the nodes by dragging from one connection point to the next.
4. **Save and Test:**
* Click ‘**Save**’ and try out your new flow to see how it works.
***
### Lesson 3b: Creating Using a Prompt
1. **Start Fresh:**
* Navigate to ‘**Agent Chains**’ and click ‘**+**’.
* Select ‘**Create from Prompt**’, name your chain, choose an industry, and pick an icon.
2. **Generate Your Chain:**
* Provide a detailed prompt describing your conversation flow.
* The system generates an initial flow based on your input.
3. **Review and Customize:**
* Edit nodes, change connections, and refine the prompts to better fit your vision.
4. **Save and Test:**
* Once you’re happy with the flow, save your work and run a test.
***
### Lesson 3c: Uploading a JSON File
1. **Prepare Your File:**
* Ensure your agent chain is saved in a JSON file. You can find an example here.
2. **Import Your Flow:**
* Create a new chain from scratch.
* Click ‘**Import**’ in the top right corner, paste your JSON content and click on ‘**Import’**.
3. **Review the Imported Flow:**
* The system will load your chain onto the canvas.
* Adjust node configurations and layout if needed.
4. **Save Your Chain:**
* Finalize and save your agent chain.
***
## Lesson 4: Configuring your Agent Chain
### Navigating the Canvas
The canvas is your workspace for designing conversations. Think of it as a whiteboard where you visually map out each step of your conversation.
* **Drag & Drop:** Move nodes by dragging them anywhere on the canvas.
* **Connect Nodes**: Hover over a node to show connection points. Click and drag from a point to draw a line, then drop onto another node's connection point to create a flow.
* **Zoom & Pan:** Pan by clicking and dragging anywhere on the canvas. To zoom, pinch using a trackpad or scroll using a mouse.
* **Organization Tools: K**eep your chain neat and easy to manage by using the tools available in the bottom left.
* ‘**Fit View**’ centers the flow and maximizes zoom out.
* ‘**Auto Layout’** arranges the nodes neatly in a top-down format.
* **Undo/redo** in the bottom bar helps manage changes.
### Configuring your Chains
* **Configure:** Link global knowledge bases, customize agent languages, time zones, LLM, transcriber, TTS, agent phone numbers and transfer number (if a transfer event node is set up).
* **Global** **Prompt:** Define common agent instructions (tone, style, etc).
* **Manage Integrations:** Connect to third-party services.
* **Manage Actions: C**reate actions to enable the agent to use the integrations. The actions can then be linked to specific nodes.
# Import Twilio number
Source: https://docs.gnani.ai/M16x_Import_Number
Description of your new file.
## Importing a number from Twilio
### What Does Importing Mean?
If you already own purchased phone numbers from Twilio, you can seamlessly import them into the Inya platform. This saves time and allows you to use your existing numbers for both testing and production calls.
### Why Import Your Twilio Numbers?
* **Efficiency:** Reuse your existing numbers without the hassle of buying new ones.
* **Central Management:** View all your active numbers in one place.
* **Enhanced Functionality:** Easily assign numbers to agents for inbound or outbound calls.
### How to Import a Twilio Number
1. **Access Active Numbers:**
* Click on the **Phone Numbers** option in the sidebar.
* Expand the section to show **Active Numbers**.
2. **Initiate the Import Process:**
* If no numbers have been imported, you’ll see an option to **Import.**
* If numbers are already present, the **Import** button will be available in the top right.
3. **Enter Your Credentials:**
* You’ll be prompted to enter your Twilio **Account SID** and **Auth Token**.
* The system validates these credentials.
* If incorrect, an error message will prompt you to re-enter the correct details.
4. **Select and Import Numbers:**
* Upon successful validation, a paginated table will display available Twilio numbers (excluding previously imported ones).
* You can select one or multiple numbers (using checkboxes and a “Select All” option).
* Confirm your selection, optionally add labels, and complete the import
*Best Practice:* Double-check the list before confirming. If any number fails to import due to duplicates or API issues, the system will notify you of the ones needing attention.
# Dashboard
Source: https://docs.gnani.ai/M21_Dashboard
The Big Picture
The **Dashboard** is your mission control, offering a high-level view of your agents and chains.
### What’s on the Dashboard?
* **Total Calls:** The overall number of calls or chats across your account.
* **Call Duration:** Average conversation length for each agent and across all agents.
* **Dialed vs. Connected Calls:** A comparison of calls made versus those successfully connected.
### How to Access the Dashboard
* Navigate to the **Dashboard** from the sidebar. This centralized view aggregates data from all agents and chains.
### Why Use the Dashboard?
* **Quick Overview:** Get immediate insights into your overall performance.
* **Performance Trends:** Monitor trends over time to spot improvements or areas for adjustment.
* **Strategic Decisions:** Use data insights to scale and optimize agent interactions.
***
# Add Your Custom Scorecard
Source: https://docs.gnani.ai/add-your
Scorecard Forms in **Aura** are structured evaluation templates used to assess the performance of human agents during customer calls.
These forms allow supervisors and quality analysts to review interactions using predefined criteria and assign scores based on the agent’s performance. Each scorecard form contains **sections and questions** that represent different aspects of the call, such as communication quality, issue handling, compliance, and customer experience.
The evaluation results help organizations monitor agent performance, maintain service quality, and generate insights through Aura’s analytics.
***
# Why Scorecard Forms are Important
Scorecard forms provide a **standardized way to evaluate agent calls**. Without structured evaluation criteria, quality monitoring becomes subjective and inconsistent.
Scorecard forms help organizations:
* Maintain consistent evaluation standards across teams
* Measure agent performance using structured scoring
* Identify coaching and training opportunities
* Ensure compliance with operational guidelines
* Generate performance analytics for agents
Since Aura focuses on **AI-powered analytics for human agents**, scorecard forms act as the foundation for quality monitoring and performance analysis.
# Creating a Scorecard Form
Follow these steps to create a new scorecard form.
## Step 1: Navigate to Scorecard Forms
From the left navigation menu:
**Scorecard Forms → Create Form**
This opens the scorecard form builder.
## Step 2: Enter Form Details
Provide the basic details for the scorecard.
### Form Name
Enter the name of the scorecard.\
Example: Customer Support QA Evaluation
### Form Description
Provide a brief description explaining the purpose of the form.\
Example: Evaluates customer support calls based on greeting, communication clarity, and issue resolution.
### Marks Configuration
Aura allows score-based evaluations. When marks are enabled:
* Each question contributes to the total score
* The system automatically calculates the **maximum marks**
# Components of a Scorecard Form
A scorecard form is composed of several configurable components that define how evaluations are structured.
***
## Sections
Sections help organise evaluation criteria into logical groups.
Examples include:
* Greeting & Introduction
* Communication Skills
* Issue Resolution
* Compliance
To create a section:
1. Click **Add Section**
2. Enter the **Section Name**
3. Define the **Total Marks**
The marks assigned to the section are **equally distributed across all questions in that section**.
Example: Section Marks: 10
Questions: 2\
Each question will carry **5 marks**.\
Users can also:
* Delete sections
* Rearrange sections
* Modify section details
***
## Questions
Each section contains one or more questions used to evaluate specific aspects of the call.
Example questions:
* Did the agent greet the customer professionally?
* Did the agent clearly understand the issue?
* Was the solution communicated properly?
Users can:
* Add multiple questions
* Delete questions
* Rearrange questions within a section
## Question Format
Each question in the scorecard follows a structured format to ensure consistent evaluation.\
\
A question includes the following elements:
### Question
The evaluation statement used to assess a specific behaviour or action during the call.\
Example: Did the agent greet the customer professionally?
***
### Response Type
Defines how the evaluator responds to the question.
Aura supports multiple response formats, including:
* **Boolean** (Yes / No)
* **Multiple Choice (Bullets)**
* **Multiple Choice (Checkbox)**
* **Comment-based responses**
* **Mixed question types**
These formats allow the scorecard to capture both **quantitative scores and qualitative feedback**.
***
### Response Options
Depending on the response type, evaluators can select from predefined options.
Example: Yes / No\
Each option can be assigned a **score value**.
Example:\
Yes → 5 marks\
No → 0 marks
## Question Options
Each question can include additional evaluation options.
### Comments
When enabled, the **Comments** option allows evaluators to provide written feedback during evaluation.
This helps reviewers explain the context behind the score and provide constructive feedback for agent improvement.
Example:\
The agent greeted the customer politely but missed confirming the order number.
***
### N/A Field
The **N/A (Not Applicable)** field allows evaluators to mark a question as not relevant to the call being reviewed.
Example:\
If the call does not involve billing, billing-related questions can be marked as **N/A**.\
This ensures agents are not penalized for scenarios that did not apply to the interaction.
***
## Fatal Section
Aura allows users to create **Fatal Sections** for critical evaluation criteria.\
Fatal sections contain questions that represent serious compliance or process violations.
Examples include:
* Security verification not performed
* Incorrect information was provided to the customer
* Mandatory compliance statement skipped
Users can add multiple questions in a single section.\
If a fatal question fails, the evaluation may result in **automatic failure regardless of the total score**.
***
## AI-Generated Sections
Aura also allows users to generate scorecard sections using **AI**.
To generate a section:
1. Click the **Generate Section** option
2. Describe the section you want to create
3. Select the **number of questions**
4. Choose the **question type**
Supported question types include:
* Multiple Choice (Bullets)
* Boolean
* Comment
* Multiple Choice (Checkbox)
* Mixed Questions
* All Types
Aura will automatically generate relevant evaluation questions based on the prompt.\
This helps teams quickly build structured scorecards without manually writing each question.
***
## Previewing the Form
Before saving the scorecard, users can review the structure using the **Preview** button.
The preview allows users to check:
* Section structure
* Question order
* Response types
* Scoring distribution
***
## Saving the Form
Once the scorecard structure is complete:
Click **Save**\
The form will be available in the **Scorecard Forms dashboard** and can be used to evaluate agent calls.
# Speech-to-Text (REST)
Source: https://docs.gnani.ai/api/STT/speech-to-text
POST /stt/v3
Quick transcription of audio clips up to 60 seconds via HTTP.
## Overview
The REST endpoint transcribes an audio file in a single synchronous HTTP request and returns the transcript immediately. It is best suited for short, pre-recorded audio clips.
| Use case | Recommended endpoint |
| --------------------------------- | ------------------------------------------------------ |
| Short clips ≤ 60 s (ideal ≤ 30 s) | **This endpoint** |
| Live microphone / real-time audio | [STT Realtime (WebSocket)](/vachana/STT/stt-websocket) |
| Large files or bulk jobs | [STT Batch](/vachana/STT/stt-batch) |
***
## Endpoint
```text theme={null}
POST https://api.vachana.ai/stt/v3
Content-Type: multipart/form-data
```
***
## Authentication
Pass your API key in the request header.
| Header | Type | Required | Description |
| -------------- | -------- | -------- | ------------------------------------------------------------------------- |
| `X-API-Key-ID` | `string` | ✅ | Your Gnani Prisma v2.5 API key. Obtain one from the Gnani APIs dashboard. |
***
## Request Parameters
All parameters are sent as `multipart/form-data` fields.
Audio file to transcribe. Supported formats: WAV, MP3, OGG, FLAC, AAC, M4A. Maximum duration: 60 seconds (ideal ≤ 30 s).
BCP-47 language code. See [Supported Languages](#supported-languages) below. Pass a comma-separated list of codes to enable auto-detection.
Forces processing with the single-language model for the specified code. Must be one of the values passed in `language_code`. Useful to improve accuracy when the audio is predominantly one language.
`verbatim` — raw spoken-form output. `transcribe` — enables Inverse Text Normalization (ITN): numbers, currency, dates, and phone numbers are written in their conventional form. See [ITN](#inverse-text-normalization-itn) below.
When `format=transcribe`, set `true` to render digits in the native script of the target language (e.g. `₹५,०००` instead of `₹5,000` for Hindi). Has no effect when `format=verbatim`. Currently supported for `hi-IN` and `en-IN` only.
***
## Response
### 200 — Success
```json theme={null}
{
"success": true,
"request_id": "req_abc123",
"timestamp": "20251226_143052.123",
"transcript": "नमस्ते, आप कैसे हैं?"
}
```
| Field | Type | Description |
| ------------ | --------- | --------------------------------------------------------------------------------------- |
| `success` | `boolean` | `true` when transcription completed without error. |
| `request_id` | `string` | Unique identifier for this request. Use it when contacting support or correlating logs. |
| `timestamp` | `string` | Server-side request timestamp in `YYYYMMDD_HHMMSS.mmm` format. |
| `transcript` | `string` | The transcribed text. Format depends on the `format` parameter. |
### Error Responses
| Status | Meaning |
| ------ | ------------------------------------------------------------------------ |
| `400` | Bad request — invalid parameters or unsupported audio format. |
| `429` | Rate limit exceeded — slow down or contact support to increase limits. |
| `500` | Internal server error — transient issue on our side; retry with backoff. |
| `503` | Service unavailable — the STT service is temporarily down. |
***
## Code Example
```bash cURL theme={null}
curl --request POST \
--url https://api.vachana.ai/stt/v3 \
--header 'Content-Type: multipart/form-data' \
--header 'X-API-Key-ID: ' \
--form audio_file='@recording.wav' \
--form language_code=hi-IN \
--form format=transcribe \
--form itn_native_numerals=true
```
```python Python SDK theme={null}
from gnani.stt import GnaniSTTClient
client = GnaniSTTClient(
organization_id="your-organization-id",
api_key="your-api-key",
user_id="your-user-id",
)
result = client.transcribe("recording.wav", language_code="hi-IN")
print(result["transcript"])
```
***
## Python SDK
The official Python SDK handles multipart construction, authentication headers, and retries automatically.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
The client requires three credentials: `organization_id`, `api_key`, and `user_id`. You can pass them directly or load them from environment variables.
```python Constructor arguments theme={null}
from gnani.stt import GnaniSTTClient
client = GnaniSTTClient(
organization_id="your-organization-id",
api_key="your-api-key",
user_id="your-user-id",
)
```
```bash Environment variables theme={null}
export GNANI_ORGANIZATION_ID="your-organization-id"
export GNANI_API_KEY="your-api-key"
export GNANI_USER_ID="your-user-id"
```
```python Environment variables (usage) theme={null}
from gnani.stt import GnaniSTTClient
# Picks up credentials from environment automatically
client = GnaniSTTClient()
```
### Transcribe Audio
```python From a file path theme={null}
result = client.transcribe("recording.wav", language_code="hi-IN")
print(result["transcript"])
```
```python From a file object theme={null}
with open("recording.wav", "rb") as f:
result = client.transcribe(f, language_code="hi-IN")
print(result["transcript"])
```
```python From raw bytes theme={null}
with open("recording.wav", "rb") as f:
audio_bytes = f.read()
result = client.transcribe(audio_bytes, language_code="hi-IN")
print(result["transcript"])
```
### Custom Request ID
Pass a `request_id` to correlate SDK calls with your own logs or support tickets.
```python theme={null}
result = client.transcribe(
"call.flac",
language_code="hi-IN",
request_id="my-trace-123",
)
```
### Error Handling
```python theme={null}
from gnani.stt import (
AuthenticationError,
InvalidAudioError,
APIError,
)
try:
result = client.transcribe("audio.wav", language_code="hi-IN")
print(result["transcript"])
except AuthenticationError:
print("Invalid credentials — check your organization_id, api_key, and user_id.")
except InvalidAudioError as e:
print(f"Bad audio file: {e}")
except APIError as e:
print(f"API error {e.status_code}: {e}")
```
***
## Supported Languages
The Gnani Prisma v2.5 API supports 10 Indian languages.
| Language | Code | Native Script | Example |
| --------- | ------- | ------------------- | ------------------------------- |
| Bengali | `bn-IN` | Bengali (বাংলা) | "আমি ভাত খাই" |
| English | `en-IN` | Latin | "I am going to the market" |
| Gujarati | `gu-IN` | Gujarati (ગુજરાતી) | "હું બજાર જાઉં છું" |
| Hindi | `hi-IN` | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
| Kannada | `kn-IN` | Kannada (ಕನ್ನಡ) | "ನಾನು ಮಾರುಕಟ್ಟೆಗೆ ಹೋಗುತ್ತೇನೆ" |
| Malayalam | `ml-IN` | Malayalam (മലയാളം) | "ഞാൻ ചന്തയിലേക്ക് പോകുന്നു" |
| Marathi | `mr-IN` | Devanagari (मराठी) | "मी बाजारात जातोय" |
| Punjabi | `pa-IN` | Gurmukhi (ਪੰਜਾਬੀ) | "ਮੈਂ ਬਾਜ਼ਾਰ ਜਾ ਰਿਹਾ ਹਾਂ" |
| Tamil | `ta-IN` | Tamil (தமிழ்) | "நான் சந்தைக்கு செல்கிறேன்" |
| Telugu | `te-IN` | Telugu (తెలుగు) | "నేను మార్కెట్కి వెళ్తున్నాను" |
For **auto-detection**, pass the full set of candidate language codes as a comma-separated value in `language_code`. For example: `en-IN,hi-IN,ta-IN,te-IN,kn-IN,ml-IN,gu-IN,mr-IN,bn-IN,pa-IN`.
***
## Inverse Text Normalization (ITN)
ITN converts the spoken-form output of the ASR engine into the conventional written form a reader expects — numbers become digits, currency gets the ₹ symbol, dates are formatted, and phone numbers are compacted — all in one pass, immediately after transcription.
**How to enable:** Set `format=transcribe` in the request body.
Currently supported for **Hindi (`hi-IN`)** and **English (`en-IN`)** only. All other languages use `verbatim` output regardless of the `format` value.
### What ITN Normalizes
#### 1 — Cardinal & Ordinal Numbers
Whole numbers and positional ranks are formatted using Indian comma grouping (groups of 2 after the first 3 digits).
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------- | -------------------- | ----------------------- |
| दो हज़ार | 2,000 | Indian comma grouping |
| पाँच लाख बीस हज़ार | 5,20,000 | Lakh-scale grouping |
| उन्नीस सौ चौरानवे | 1,994 | Hundred-base year form |
| five lakh | 5,00,000 | English lakh convention |
| पहला / twenty first | 1st / 21st | Ordinal suffix |
#### 2 — Currency & Money
All Indian currency expressions — including paise fractions and lakh/crore scales — are formatted with the ₹ symbol and Indian comma grouping.
| Spoken input (ASR) | Written output (ITN) | Rule |
| --------------------------- | -------------------- | ---------------------- |
| पाँच सौ रुपये | ₹500 | ₹ + amount |
| तीन रुपये पचास पैसे | ₹3.50 | ₹ + rupees.paise |
| दस लाख रुपये | ₹10,00,000 | ₹ + lakh grouping |
| I need five thousand rupees | ₹5,000 | English India pipeline |
#### 3 — Dates
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------ | -------------------- | ----------------------- |
| बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 | DD Month YYYY (hi) |
| fifteenth january twenty twenty five | 15th January 2025 | Ordinal Month YYYY (en) |
#### 4 — Times
Indian time-of-day words (सुबह, दोपहर, शाम, रात) automatically map to 24-hour HH:MM output.
| Spoken input (ASR) | Written output (ITN) | Rule |
| -------------------------------------- | ---------------------------- | ----------------------- |
| सुबह पाँच बजे | सुबह 05:00 | सुबह = AM |
| शाम पाँच बजे | शाम 17:00 | शाम = evening (16–20 h) |
| रात के दस बजे | रात 22:00 | रात = night (20–24 h) |
| meeting at five fifteen in the evening | meeting 17:15 in the evening | en — 24-hour |
#### 5 — Phone Numbers & PIN Codes
Digit streams are concatenated into compact numeric strings. 10-digit streams → mobile number; 6-digit streams → PIN. Repeat prefixes (double/डबल, triple/ट्रिपल) are expanded.
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------- | -------------------- | ------------------- |
| नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 | 10 digits → phone |
| एक एक शून्य शून्य शून्य एक | 110001 | 6 digits → PIN |
| डबल आठ नौ शून्य एक दो तीन चार पाँच छह | 8890123456 | double prefix |
| one two three four five six | 123456 | English digit words |
#### 6 — Mixed & Code-Mixed Utterances
A single sentence may contain multiple entity types or blend Hindi and English. ITN handles all in one pass, normalizing each entity independently.
| Spoken input (ASR) | Written output (ITN) |
| ----------------------------------------------------- | --------------------------------- |
| कल थ्री फिफ्टी पीएम को पाँच सौ रुपये transfer करना है | कल 15:50 को ₹500 transfer करना है |
### Native Script Digits — `itn_native_numerals`
By default, ITN outputs Western Arabic digits (0–9) regardless of language. Set `itn_native_numerals=true` to render digits in the native script of the target language.
| Language | Spoken input | `false` (default) | `true` — native script |
| --------------- | -------------------- | ----------------- | -------------------------- |
| Hindi `hi-IN` | पाँच हज़ार रुपये | ₹5,000 | ₹५,००० |
| English `en-IN` | five thousand rupees | ₹5,000 | ₹5,000 (Latin — no change) |
### What ITN Does Not Change
ITN intentionally preserves idiomatic and ambiguous phrases to avoid incorrect normalization.
* **दो तीन** (meaning *a few*) stays as text, not `2` or `3`
* **कर दो / ले दो** (imperative verbs) are kept as words, not treated as cardinal 2
If a word or phrase is unchanged in the output, treat it as a failure only when the input was unambiguously a numeric entity.
# Speech-to-Text (Batch)
Source: https://docs.gnani.ai/api/STT/stt-batch
POST /stt/v3/batch/submit
Asynchronous transcription of long or multiple audio files via HTTP.
## Overview
Submit one or more audio files for transcription and receive a `job_id` immediately. Poll the status endpoint on a fixed interval until the job completes and transcripts are available. Ideal for long recordings, bulk files, or offline pipelines where you do not need a live response. For real-time transcription, see [STT Realtime](/vachana/STT/stt-websocket). For short clips under 60 seconds, see [STT REST](/vachana/STT/speech-to-text).
## Endpoints
| Operation | Method | URL |
| ---------------- | ------ | ----------------------------------------------------- |
| Submit Job | POST | `https://api.vachana.ai/stt/v3/batch/submit` |
| Check Job Status | GET | `https://api.vachana.ai/stt/v3/batch/status/{job_id}` |
## Limits & Specifications
| Item | Limit |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| Max audio duration | Less than **1 hour** per file |
| Max files per request | **10 files** per API call |
| Max total payload size | **80 MB** across all files and form fields combined |
| Minimum poll interval | **60 seconds** between status calls for the same `job_id` |
| Speaker diarization. | This API supports **at most 2 speakers** per file (two-party diarization). Scenarios with more than two distinct speakers are not supported |
### Supported Audio Formats
`AAC` · `WAV` · `FLAC` · `ALAC` · `OGG (Vorbis)` · `Opus`
Use standard file extensions and MIME types (e.g. `.m4a` for AAC, `.wav`, `.flac`, `.ogg`).
## Authentication
Send these headers on **every request** both submit and status calls.
| Header | Required | Description |
| ------------------ | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| `X-API-Key-ID` | Yes | Your API key. Required for all requests. |
| `X-API-Request-ID` | No | A unique trace ID (e.g. UUID) you assign. Used to correlate your logs with platform logs or support. If omitted, the platform may generate one. |
Do **not** set `Content-Type: application/json` on the submit request. Use `multipart/form-data`. curl sets the correct boundary automatically when you use `-F` / `--form`.
***
## Submit a Transcription Job
### `POST /stt/v3/batch/submit`
Upload audio files and kick off an asynchronous transcription job. The response returns a `job_id` immediately. the files are not yet transcribed at this point.
#### Request — Form Fields
#### Supported Language Codes
| Language | Code | Native Script | Example Text |
| --------- | ------- | ------------------- | ------------------------------- |
| Bengali | `bn-IN` | Bengali (বাংলা) | "আমি ভাত খাই" |
| English | `en-IN` | Latin | "I am going to the market" |
| Gujarati | `gu-IN` | Gujarati (ગુજરાતી) | "હું બજાર જાઉં છું" |
| Hindi | `hi-IN` | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
| Kannada | `kn-IN` | Kannada (ಕನ್ನಡ) | "ನಾನು ಮಾರುಕಟ್ಟೆಗೆ ಹೋಗುತ್ತೇನೆ" |
| Malayalam | `ml-IN` | Malayalam (മലയാളം) | "ഞാൻ ചന്തയിലേക്ക് പോകുന്നു" |
| Marathi | `mr-IN` | Devanagari (मराठी) | "मी बाजारात जातोय" |
| Punjabi | `pa-IN` | Gurmukhi (ਪੰਜਾਬੀ) | "ਮੈਂ ਬਾਜ਼ਾਰ ਜਾ ਰਿਹਾ ਹਾਂ" |
| Tamil | `ta-IN` | Tamil (தமிழ்) | "நான் சந்தைக்கு செல்கிறேன்" |
| Telugu | `te-IN` | Telugu (తెలుగు) | "నేను మార్కెట్కి వెళ్తున్నాను" |
#### Example — curl
```bash theme={null}
curl --location --request POST 'https://api.vachana.ai/stt/v3/batch/submit' \
--header 'X-API-Key-ID: ' \
--header 'X-API-Request-ID: 550e8400-e29b-41d4-a716-446655440000' \
--form 'language_code=hi-IN' \
--form 'is_multi_channel=false' \
--form 'format=transcribe' \
--form 'audio_files=@"/path/to/first.wav"' \
--form 'audio_files=@"/path/to/second.wav"'
```
#### Response — `200 OK`
```json theme={null}
{
"job_id": "batch_7f3a92c1d4e8",
"status": "submitted",
"file_count": 2,
"message": "Job accepted. Poll the status endpoint every 60 seconds for results."
}
```
| Field | Type | Description |
| ------------ | ------- | -------------------------------------------------- |
| `job_id` | string | Identifier for this job. Use it in the status URL. |
| `status` | string | Initial value is always `submitted`. |
| `file_count` | integer | Number of files accepted into the job. |
| `message` | string | Short confirmation with polling instructions. |
#### Errors
| HTTP Status | When |
| ----------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `400` | No files uploaded, empty file, more than 10 files, payload over 80 MB, unsupported format, or other client-side validation failure. |
| `500` | Server error. |
***
## Check Job Status
### `GET /stt/v3/batch/status/{job_id}`
Poll this endpoint to check progress and retrieve transcription results once the job finishes. Call this **once every 60 seconds** per `job_id`. do not poll more frequently.
#### Path Parameter
| Parameter | Required | Description |
| --------- | -------- | ----------------------------------------------- |
| `job_id` | Yes | The `job_id` returned from the Submit response. |
#### Example — curl
```bash theme={null}
curl --location --request GET 'https://api.vachana.ai/stt/v3/batch/status/{job_id}' \
--header 'X-API-Key-ID: ' \
--header 'X-API-Request-ID: 550e8400-e29b-41d4-a716-446655440000'
```
#### Response — `200 OK`
```json theme={null}
{
"job_id": "batch_7f3a92c1d4e8",
"status": "completed",
"total_files": 2,
"completed_files": 2,
"failed_files": 0,
"overall_progress": 100,
"error": null,
"results": [
{
"filename": "first.wav",
"status": "completed",
"full_transcript": "नमस्ते, आप कैसे हैं?",
"total_duration": 45.3,
"error": null,
"segments": [
{
"segment_id": 0,
"start_time": 0.0,
"end_time": 3.2,
"text": "नमस्ते, आप कैसे हैं?",
"speaker_id": 1,
"language_detected": "hi-IN"
}
]
}
]
}
```
#### Job-Level Response Fields
| Field | Type | Description |
| ------------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| `job_id` | string | Job identifier. |
| `status` | string | `submitted` — accepted or in progress. `processing` — actively transcribing. `completed` — done. `failed` — job-level failure. |
| `total_files` | integer | Total number of files in the job. |
| `completed_files` | integer | Files finished successfully. Meaningful only when the job has reached a final state. |
| `failed_files` | integer | Files that failed. Meaningful only when the job has reached a final state. |
| `overall_progress` | integer | Approximate progress from `0` to `100` while the job is running. |
| `results` | array or `null` | Per-file results. `null` while the job is `submitted` or `processing`. |
| `error` | string or `null` | Top-level error message for the job, if any. |
#### Per-File Result Fields — `results[]`
| Field | Type | Description |
| ----------------- | ---------------- | ------------------------------------------------- |
| `filename` | string | Original file name as submitted. |
| `full_transcript` | string | Complete transcribed text for the file. |
| `segments` | array | Time-aligned transcript segments (see below). |
| `total_duration` | number | Audio duration in seconds. |
| `status` | string | `completed` or `failed` for this individual file. |
| `error` | string or `null` | Error message for this file if it failed. |
#### Per-Segment Fields — `results[].segments[]`
| Field | Type | Description |
| ------------------- | ------- | ------------------------------------------------------ |
| `segment_id` | integer | Segment index (zero-based). |
| `start_time` | number | Segment start time in seconds. |
| `end_time` | number | Segment end time in seconds. |
| `text` | string | Transcribed text for this segment. |
| `speaker_id` | integer | Speaker identifier. Populated for multi-channel audio. |
| `language_detected` | string | BCP-47 code of the detected language for this segment. |
#### Errors
| HTTP Status | When |
| ----------- | ------------------------------------------------------------------ |
| `404` | `job_id` not found — unknown ID or the job is no longer available. |
| `500` | Server error. |
***
## Inverse Text Normalization (ITN)
When `format=transcribe` is passed in the form body, ITN runs on every file's transcript after recognition — converting spoken-form numbers, currency, dates, times, and phone numbers into the compact written form a reader expects.
ITN is currently supported for **Hindi (`hi-IN`)** and **English (`en-IN`)** only. Enabling ITN for other languages has no effect — transcripts are returned as-is.
#### What ITN Normalizes
ITN recognizes six categories of spoken expressions. Every matching span is transformed; all other words pass through unchanged.
Whole numbers and positional ranks are formatted using Indian comma grouping (groups of 2 after the first 3 digits).
| Spoken input (ASR) | Written output (ITN) | Format rule |
| ------------------- | -------------------- | ----------------------- |
| दो हज़ार | 2,000 | Indian comma grouping |
| पाँच लाख बीस हज़ार | 5,20,000 | Lakh-scale grouping |
| five lakh | 5,00,000 | English lakh convention |
| पहला / twenty first | 1st / 21st | Ordinal suffix |
All Indian currency expressions — including paise fractions and lakh/crore scales — are formatted with the ₹ symbol and Indian comma grouping.
| Spoken input (ASR) | Written output (ITN) | Format rule |
| --------------------------- | -------------------- | ---------------------- |
| पाँच सौ रुपये | ₹500 | ₹ + amount |
| तीन रुपये पचास पैसे | ₹3.50 | ₹ + rupees.paise |
| I need five thousand rupees | ₹5,000 | English India pipeline |
| pay do lakh rupees | ₹2,00,000 | Code-mixed en/hi |
| Spoken input (ASR) | Written output (ITN) | Format rule |
| ------------------------------------ | -------------------- | ----------------------- |
| बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 | DD Month YYYY (hi) |
| fifteenth january twenty twenty five | 15th January 2025 | Ordinal Month YYYY (en) |
Indian time-of-day words (सुबह, दोपहर, शाम, रात) automatically map to 24-hour HH:MM output.
| Spoken input (ASR) | Written output (ITN) | Format rule |
| -------------------------------------- | ---------------------------- | ---------------------- |
| सुबह पाँच बजे | सुबह 05:00 | सुबह = AM |
| शाम पाँच बजे | शाम 17:00 | शाम = evening (16–20h) |
| रात के दस बजे | रात 22:00 | रात = night (20–24h) |
| meeting at five fifteen in the evening | meeting 17:15 in the evening | en — 24-hour |
10-digit streams → mobile number; 6-digit streams → PIN. Repeat prefixes (double/डबल) are expanded.
| Spoken input (ASR) | Written output (ITN) | Format rule |
| ------------------------------------- | -------------------- | ------------------- |
| नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 | 10 digits → phone |
| एक एक शून्य शून्य शून्य एक | 110001 | 6 digits → PIN |
| one two three four five six | 123456 | English digit words |
A single file may contain segments with multiple entity types or blend Hindi and English. ITN normalizes each entity independently in one pass.
| Spoken input (ASR) | Written output (ITN) |
| ----------------------------------------------------- | --------------------------------- |
| कल थ्री फिफ्टी पीएम को पाँच सौ रुपये transfer करना है | कल 15:50 को ₹500 transfer करना है |
| pay do lakh rupees by fifteenth march | pay ₹2,00,000 by 15th March |
#### Native Script Digits — `itn_native_numerals`
By default, ITN outputs Western Arabic digits (0–9) regardless of language. When `format=transcribe` is set, you can additionally pass `itn_native_numerals=true` to render digits in the native script of the target language.
| Language | Spoken input | `false` (default) | `true` — native script |
| --------------- | -------------------- | ----------------- | -------------------------- |
| Hindi `hi-IN` | पाँच हज़ार रुपये | ₹5,000 | ₹५,००० |
| English `en-IN` | five thousand rupees | ₹5,000 | ₹5,000 (Latin — no change) |
English always outputs Western Arabic digits. `itn_native_numerals=true` has no effect for `en-IN`.
#### What ITN Does Not Change
ITN intentionally preserves idiomatic and ambiguous phrases.
* **दो तीन** (meaning *a few*) stays as text, not `2` or `3`
* **कर दो / ले दो** (imperative verbs) are kept as words, not treated as cardinal 2
***
## Flow Summary
1. **Submit** — `POST https://api.vachana.ai/stt/v3/batch/submit` with `X-API-Key-ID`, optional `X-API-Request-ID`, and form fields `audio_files`, `language_code`, and optionally `is_multi_channel`, `format`, and `itn_native_numerals`.
2. **Save the `job_id`** from the submit response.
3. **Poll** — `GET https://api.vachana.ai/stt/v3/batch/status/{job_id}` (same auth headers) **every 60 seconds** until `status` is `completed` or `failed` and `results` is populated.
# Speech-to-Text (Realtime)
Source: https://docs.gnani.ai/api/STT/stt-websocket
Real-time speech-to-text over a persistent WebSocket connection.
## Overview
Stream raw PCM audio frames and receive transcript segments as speech is detected. The server uses Voice Activity Detection (VAD) to identify speech boundaries and returns a transcript for each segment.
| Use case | Recommended endpoint |
| ---------------------------------------------- | --------------------------------------- |
| Live microphone / phone call / real-time audio | **This endpoint** |
| Short pre-recorded clips ≤ 60 s | [STT REST](/vachana/STT/speech-to-text) |
| Large files or bulk jobs | [STT Batch](/vachana/STT/stt-batch) |
***
## Endpoint
```text theme={null}
WSS wss://api.vachana.ai/stt/v3/stream
```
***
## Connection Headers
All configuration is passed as WebSocket upgrade headers at connection time. Headers cannot be changed mid-session — reconnect with new headers to change settings.
| Header | Required | Default | Description |
| --------------------- | -------- | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| `x-api-key-id` | ✅ | — | Your Gnani API key. |
| `lang_code` | ✅ | `en-IN` | BCP-47 language code for transcription. See [Supported Languages](#supported-languages). |
| `x-sample-rate` | ❌ | `16000` | Sample rate of the audio stream in Hz. Accepted values: `8000`, `16000`, `44100`, `48000`. Must match the actual sample rate of your audio source. |
| `x-format` | ❌ | `verbatim` | `verbatim` — raw spoken-form output. `transcribe` — enables Inverse Text Normalization (ITN). See [ITN](#inverse-text-normalization-itn). |
| `itn_native_numerals` | ❌ | `false` | When `x-format=transcribe`, set `true` to render digits in the native script of the target language (e.g. `₹५,०००` instead of `₹5,000` for Hindi). |
**Choosing the right sample rate:**
| Value | When to use |
| ------- | ------------------------------------------------- |
| `48000` | Browser `getUserMedia` default; Mac microphone |
| `44100` | Mac microphone alternate; CD-quality audio |
| `16000` | Wideband telephony; sent as-is with no resampling |
| `8000` | Narrowband telephony (legacy PSTN / VoIP) |
***
## Connection Flow
A WebSocket session follows a strict sequence:
1. **Client connects** — opens a WebSocket to `/stt/v3/stream` with all required headers.
2. **Server confirms** — immediately sends a `connected` message echoing the active configuration.
3. **Client streams audio** — continuously sends binary frames of raw PCM audio at a steady real-time cadence.
4. **Server detects speech** — VAD identifies end-of-speech boundaries and emits a `processing` message to acknowledge that a segment was captured.
5. **Server returns transcript** — sends a `transcript` message with the transcribed text, segment metadata, and latency.
6. **Either side closes** — client or server may close the connection at any time.
The `processing` message is a low-latency signal that audio was captured and transcription has begun. Expect a `transcript` message shortly after.
***
## Audio Format & Sending Audio
All audio must be sent as **raw PCM binary frames** over the WebSocket. No container format (WAV, MP3, etc.) is accepted mid-stream.
### PCM Specification
| Property | 16 kHz | 8 kHz |
| ------------------- | --------------------------------------- | --------------------------------------- |
| Encoding | PCM signed 16-bit little-endian | PCM signed 16-bit little-endian |
| Sample Rate | 16,000 Hz | 8,000 Hz |
| Channels | 1 (mono) | 1 (mono) |
| Samples per chunk | 512 | 512 |
| **Bytes per frame** | **1,024 bytes** (512 samples × 2 bytes) | **1,024 bytes** (512 samples × 2 bytes) |
| Frame duration | 32 ms | 64 ms |
### Sending Rules
* Each binary frame must be **exactly 1,024 bytes**.
* Frames must be sent at **real-time cadence** — one frame every 32 ms (16 kHz) or 64 ms (8 kHz). Do not buffer and burst; this degrades VAD accuracy.
* For `44100` and `48000` Hz sources, the server resamples internally — still send 1,024-byte frames at the appropriate cadence.
***
## Server Messages
The server sends JSON text frames. All messages share a `type` discriminator field and an ISO-8601 `timestamp`.
### `connected`
Sent once immediately after the WebSocket handshake succeeds.
```json theme={null}
{
"type": "connected",
"message": "STT service ready — VAD service connected",
"timestamp": "2024-01-15T10:30:00.000Z",
"config": {
"sample_rate": 16000,
"chunk_size": 512
}
}
```
| Field | Type | Description |
| -------------------- | --------- | ------------------------------------------------------ |
| `type` | `string` | Always `"connected"`. |
| `message` | `string` | Human-readable status string. |
| `timestamp` | `string` | ISO-8601 server timestamp. |
| `config.sample_rate` | `integer` | Active sample rate in Hz, echoed from `x-sample-rate`. |
| `config.chunk_size` | `integer` | Expected chunk size in samples (always 512). |
### `processing`
Emitted when VAD detects the end of a speech segment and transcription has begun. Use this as a low-latency acknowledgment that audio was captured.
```json theme={null}
{
"type": "processing",
"timestamp": "2024-01-15T10:30:05.123Z"
}
```
| Field | Type | Description |
| ----------- | -------- | ------------------------------------------------ |
| `type` | `string` | Always `"processing"`. |
| `timestamp` | `string` | ISO-8601 timestamp when speech-end was detected. |
### `transcript`
Contains the transcribed text for a completed speech segment.
```json theme={null}
{
"type": "transcript",
"timestamp": "2024-01-15T10:30:05.987Z",
"text": "Hello, how are you today?",
"audio_duration_ms": 2340,
"segment_id": "",
"segment_index": 0,
"latency": 320
}
```
| Field | Type | Description |
| ------------------- | --------- | ---------------------------------------------------------------------------------------- |
| `type` | `string` | Always `"transcript"`. |
| `timestamp` | `string` | ISO-8601 timestamp when the transcript was emitted. |
| `text` | `string` | Transcribed text. Format depends on the `x-format` header. |
| `audio_duration_ms` | `integer` | Duration of the captured speech segment in milliseconds. |
| `segment_id` | `string` | Unique identifier for this speech segment. Use for deduplication or support correlation. |
| `segment_index` | `integer` | Sequential index of this segment within the session, starting at `0`. |
| `latency` | `integer` | Time in milliseconds from end-of-speech detection to transcript delivery. |
### `error`
Sent when the server encounters a recoverable or fatal error. The connection may remain open after a recoverable error.
| Field | Type | Description |
| ----------- | -------- | ---------------------------------------- |
| `type` | `string` | Always `"error"`. |
| `timestamp` | `string` | ISO-8601 timestamp of the error. |
| `message` | `string` | Human-readable description of the error. |
```json theme={null}
{
"type": "error",
"timestamp": "2024-01-15T10:30:10.000Z",
"message": "STT engine failed to initialize"
}
```
***
## Python SDK
The official Python SDK wraps the WebSocket connection, audio pacing, and event parsing into a clean async interface.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
The streaming client requires your API key and language code.
```python Constructor argument theme={null}
from gnani.stt import GnaniSTTStreamClient
stream = GnaniSTTStreamClient(
api_key="your-api-key",
language_code="hi-IN",
)
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.stt import GnaniSTTStreamClient
# Picks up GNANI_API_KEY from environment automatically
stream = GnaniSTTStreamClient(language_code="hi-IN")
```
### Stream Audio from a File
Use the async context manager and the `stream_audio` helper. It handles real-time pacing automatically so frames are sent at the correct cadence for VAD.
```python theme={null}
import asyncio
from gnani.stt import GnaniSTTStreamClient
async def main():
async with GnaniSTTStreamClient(
api_key="your-api-key",
language_code="hi-IN",
sample_rate=16000,
) as stream:
with open("audio.pcm", "rb") as f:
transcripts = await stream.stream_audio(
f,
on_transcript=lambda t: print(f"Transcript: {t.text}"),
on_processing=lambda p: print("Processing..."),
realtime_pace=True, # sends frames at real-time cadence
)
print(f"Total segments: {len(transcripts)}")
asyncio.run(main())
```
### Iterate Over Events Manually
For lower-level control — handling each event type differently or interleaving sending and receiving — iterate over the stream directly.
```python theme={null}
import asyncio
from gnani.stt import GnaniSTTStreamClient, StreamTranscriptEvent, StreamProcessingEvent
async def main():
async with GnaniSTTStreamClient(
api_key="your-api-key",
language_code="hi-IN",
) as stream:
with open("audio.pcm", "rb") as f:
while chunk := f.read(1024):
await stream.send_audio(chunk)
await asyncio.sleep(0.032) # 32 ms per frame at 16 kHz
async for event in stream:
if isinstance(event, StreamTranscriptEvent):
print(f"[Segment {event.segment_index}] {event.text}")
print(f" Duration: {event.audio_duration_ms} ms Latency: {event.latency} ms")
elif isinstance(event, StreamProcessingEvent):
print("Processing speech...")
asyncio.run(main())
```
### Using 8 kHz Audio (Telephony)
```python theme={null}
stream = GnaniSTTStreamClient(
api_key="your-api-key",
language_code="en-IN",
sample_rate=8000,
)
```
### SDK Event Types
All events are typed dataclasses with a `raw` field containing the full server JSON.
| Event class | Key fields | Description |
| ----------------------- | ------------------------------------------------------- | -------------------------------------------------------------------- |
| `StreamConnectedEvent` | `message`, `sample_rate`, `chunk_size` | Sent once after the WebSocket handshake. Confirms the active config. |
| `StreamProcessingEvent` | `timestamp` | VAD detected end-of-speech; transcription has started. |
| `StreamTranscriptEvent` | `text`, `segment_index`, `audio_duration_ms`, `latency` | Completed transcript for a speech segment. |
| `StreamErrorEvent` | `message`, `timestamp` | Server-side error, recoverable or fatal. |
### Error Handling
```python theme={null}
from gnani.stt import (
StreamConnectionError, # Could not establish the WebSocket connection
StreamClosedError, # Attempted to send on an already-closed stream
StreamError, # Server returned an error message mid-session
)
try:
async with GnaniSTTStreamClient(api_key="your-api-key") as stream:
await stream.send_audio(chunk)
except StreamConnectionError as e:
print(f"Could not connect: {e}")
except StreamClosedError as e:
print(f"Stream was already closed: {e}")
except StreamError as e:
print(f"Server error: {e.message} (at {e.timestamp})")
```
***
## Supported Languages
| Language | Code | Native Script | Example |
| --------- | ------- | ------------------- | ------------------------------- |
| Bengali | `bn-IN` | Bengali (বাংলা) | "আমি ভাত খাই" |
| English | `en-IN` | Latin | "I am going to the market" |
| Gujarati | `gu-IN` | Gujarati (ગુજરાતી) | "હું બજાર જાઉં છું" |
| Hindi | `hi-IN` | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
| Kannada | `kn-IN` | Kannada (ಕನ್ನಡ) | "ನಾನು ಮಾರುಕಟ್ಟೆಗೆ ಹೋಗುತ್ತೇನೆ" |
| Malayalam | `ml-IN` | Malayalam (മലയാളം) | "ഞാൻ ചന്തയിലേക്ക് പോകുന്നു" |
| Marathi | `mr-IN` | Devanagari (मराठी) | "मी बाजारात जातोय" |
| Punjabi | `pa-IN` | Gurmukhi (ਪੰਜਾਬੀ) | "ਮੈਂ ਬਾਜ਼ਾਰ ਜਾ ਰਿਹਾ ਹਾਂ" |
| Tamil | `ta-IN` | Tamil (தமிழ்) | "நான் சந்தைக்கு செல்கிறேன்" |
| Telugu | `te-IN` | Telugu (తెలుగు) | "నేను మార్కెట్కి వెళ్తున్నాను" |
For **auto-detection**, pass all desired language codes comma-separated in the `lang_code` header. For example: `en-IN,hi-IN,ta-IN,te-IN,kn-IN,ml-IN,gu-IN,mr-IN,bn-IN,pa-IN`.
***
## Inverse Text Normalization (ITN)
When `x-format: transcribe` is set, ITN runs on every transcript segment immediately after recognition — converting spoken-form numbers, currency, dates, times, and phone numbers into the compact written form a reader expects.
Currently supported for **Hindi (`hi-IN`)** and **English (`en-IN`)** only. Enabling ITN for other languages has no effect; transcripts are returned verbatim.
### What ITN Normalizes
#### 1 — Cardinal & Ordinal Numbers
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------- | -------------------- | ----------------------- |
| दो हज़ार | 2,000 | Indian comma grouping |
| पाँच लाख बीस हज़ार | 5,20,000 | Lakh-scale grouping |
| five lakh | 5,00,000 | English lakh convention |
| पहला / twenty first | 1st / 21st | Ordinal suffix |
#### 2 — Currency & Money
| Spoken input (ASR) | Written output (ITN) | Rule |
| --------------------------- | -------------------- | ---------------------- |
| पाँच सौ रुपये | ₹500 | ₹ + amount |
| तीन रुपये पचास पैसे | ₹3.50 | ₹ + rupees.paise |
| I need five thousand rupees | ₹5,000 | English India pipeline |
| pay do lakh rupees | ₹2,00,000 | Code-mixed en/hi |
#### 3 — Dates
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------ | -------------------- | ----------------------- |
| बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 | DD Month YYYY (hi) |
| fifteenth january twenty twenty five | 15th January 2025 | Ordinal Month YYYY (en) |
#### 4 — Times
Indian time-of-day words (सुबह, दोपहर, शाम, रात) automatically map to 24-hour HH:MM output.
| Spoken input (ASR) | Written output (ITN) | Rule |
| -------------------------------------- | ---------------------------- | ----------------------- |
| सुबह पाँच बजे | सुबह 05:00 | सुबह = AM |
| शाम पाँच बजे | शाम 17:00 | शाम = evening (16–20 h) |
| रात के दस बजे | रात 22:00 | रात = night (20–24 h) |
| meeting at five fifteen in the evening | meeting 17:15 in the evening | en — 24-hour |
#### 5 — Phone Numbers & PIN Codes
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------- | -------------------- | ------------------- |
| नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 | 10 digits → phone |
| एक एक शून्य शून्य शून्य एक | 110001 | 6 digits → PIN |
| one two three four five six | 123456 | English digit words |
#### 6 — Mixed & Code-Mixed Utterances
| Spoken input (ASR) | Written output (ITN) |
| ----------------------------------------------------- | --------------------------------- |
| कल थ्री फिफ्टी पीएम को पाँच सौ रुपये transfer करना है | कल 15:50 को ₹500 transfer करना है |
| pay do lakh rupees by fifteenth march | pay ₹2,00,000 by 15th March |
### Native Script Digits — `itn_native_numerals`
By default, ITN outputs Western Arabic digits (0–9). Set `itn_native_numerals: true` in the connection headers to render digits in the native script of the target language.
| Language | Spoken input | `false` (default) | `true` — native script |
| --------------- | -------------------- | ----------------- | -------------------------- |
| Hindi `hi-IN` | पाँच हज़ार रुपये | ₹5,000 | ₹५,००० |
| English `en-IN` | five thousand rupees | ₹5,000 | ₹5,000 (Latin — no change) |
### What ITN Does Not Change
ITN intentionally preserves idiomatic and ambiguous phrases to avoid incorrect normalization.
* **दो तीन** (meaning *a few*) stays as text, not `2` or `3`
* **कर दो / ले दो** (imperative verbs) are kept as words, not treated as cardinal 2
***
# Text-to-Speech (REST)
Source: https://docs.gnani.ai/api/TTS/tts-inference
POST /api/v1/tts/inference
Synchronous text-to-speech with full audio returned in one response.
**Currently in beta.** You're on the priority waitlist and among the first to get access.
## Overview
Get the complete synthesized audio in one response. Best for downloads or batch processing. For streaming playback, see [TTS Streaming](/vachana/TTS/tts-sse) or [TTS Realtime](/vachana/TTS/tts-websocket).
Passing numbers, IDs, dates, or currency as raw strings causes mispronunciations. See the [Input Formatting Guide](/vachana/TTS/tts-input-formating) for correct formatting of phone numbers, account numbers, PINs, Aadhaar, vehicle registration numbers, GSTIN, currency, and more.
***
## Available Voices
| Voice | Gender | Description |
| ------- | ------ | ------------------------ |
| Pranav | Male | Bold, Trustworthy |
| Kaveri | Female | Confident, Bright |
| Shubhra | Female | Gentle, Expressive |
| Deepak | Male | Grounded, Conversational |
***
## Python SDK
The official Python SDK lets you synthesize speech in one line, without constructing JSON payloads or handling binary audio responses manually.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
The TTS client requires only your API key.
```python Constructor argument theme={null}
from gnani.tts import GnaniTTSClient
client = GnaniTTSClient(api_key="your-api-key")
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.tts import GnaniTTSClient
# Picks up GNANI_API_KEY automatically
client = GnaniTTSClient()
```
### Synthesize Speech
The `synthesize` method returns the complete audio as bytes, which you can write to a file or pass directly to an audio player.
```python theme={null}
from gnani.tts import GnaniTTSClient
client = GnaniTTSClient(api_key="your-api-key")
audio = client.synthesize(
"नमस्ते, आप कैसे हैं?",
voice="sia",
)
with open("output.wav", "wb") as f:
f.write(audio)
```
### Custom Audio Config
Control the sample rate, encoding, and container format of the output audio.
```python theme={null}
from gnani.tts import GnaniTTSClient, AudioConfig
client = GnaniTTSClient(api_key="your-api-key")
audio = client.synthesize(
"यह एक टेस्ट है",
voice="raju",
audio_config=AudioConfig(
sample_rate=44100,
encoding="linear_pcm",
container="wav",
),
)
with open("output.wav", "wb") as f:
f.write(audio)
```
### List Available Voices
```python theme={null}
from gnani.tts import GnaniTTSClient
voices = GnaniTTSClient.supported_voices()
print(voices)
```
## Supported Languages
The Gnani Timbre v2.0 API supports 2 languages.
| Language | Native Script | Example |
| -------- | ------------------- | -------------------------- |
| English | Latin | "I am going to the market" |
| Hindi | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
# Input Formatting Guide
Source: https://docs.gnani.ai/api/TTS/tts-input-formating
How to format numbers, IDs, currency, and dates in the text field for correct TTS pronunciation.
## Overview
The TTS engine is sensitive to how structured data is written in the `text` field. Passing raw strings like `9876543210` or `Rs. 1,00,000` causes mispronunciations or garbled output.
**Core principle:**
* **Read digit-by-digit** (IDs, PINs, account numbers, phone numbers) → separate every character with a single space: `9 8 7 6 5 4 3 2 1 0`
* **Read as a quantity** (prices, floor numbers, years) → pass as-is: `Rs 1,00,000`
***
## 1. Phone Numbers
Separate every digit with a single space. Never pass a raw digit string — the engine reads it as a large cardinal number.
| Scenario | Raw Input | Correct TTS Input | Expected Speech |
| ------------------------ | --------------- | --------------------------- | ----------------------------------------------------- |
| Mobile (India, 10-digit) | `9876543210` | `9 8 7 6 5 4 3 2 1 0` | nine eight seven six five four three two one zero |
| Landline with STD code | `01122334455` | `0 1 1 2 2 3 3 4 4 5 5` | zero one one two two three three four four five five |
| International (+91) | `+919876543210` | `+ 9 1 9 8 7 6 5 4 3 2 1 0` | plus nine one nine eight seven six... |
| Toll-free | `18001234567` | `1 8 0 0 1 2 3 4 5 6 7` | one eight zero zero one two three four five six seven |
Do **not** pass `9876543210`, `987-654-3210`, or `(987) 654-3210` — all cause the engine to misread.
***
## 2. Account Numbers
Every character — digit or letter — must be separated by a single space. Account numbers may be alphanumeric; each character is read individually.
| Scenario | Raw Input | Correct TTS Input | Expected Speech |
| ------------------------ | -------------- | ------------------------- | -------------------------------------------------- |
| Bank account (numeric) | `012345678901` | `0 1 2 3 4 5 6 7 8 9 0 1` | zero one two three four five... |
| Alphanumeric account | `AB123456789` | `A B 1 2 3 4 5 6 7 8 9` | A B one two three four five six seven eight nine |
| Loan account | `LN2024001234` | `L N 2 0 2 4 0 0 1 2 3 4` | L N two zero two four zero zero one two three four |
| Masked (partial visible) | `XXXX1234` | `X X X X 1 2 3 4` | X X X X one two three four |
***
## 3. PIN Numbers
Each digit must be separated by a single space. The digit `0` is always spoken as **"zero"**, never **"oh"**. Leading zeros must never be dropped.
| Scenario | Raw Input | Correct TTS Input | Expected Speech |
| ------------------- | --------- | ----------------- | ----------------------------------------- |
| 4-digit ATM PIN | `4821` | `4 8 2 1` | four eight two one |
| 6-digit OTP/PIN | `739201` | `7 3 9 2 0 1` | seven three nine two zero one |
| PIN starting with 0 | `0492` | `0 4 9 2` | zero four nine two *(not: four nine two)* |
| 6-digit postal PIN | `560001` | `5 6 0 0 0 1` | five six zero zero zero one |
***
## 4. Addresses
Street and house numbers are read as quantities (not digit-by-digit). Postal PIN codes inside addresses must be digit-by-digit. Road abbreviations (MG, NH) must be letter-spaced. Replace `/` with the word `slash`.
| Component | Raw Input | Correct TTS Input | Expected Speech |
| --------------------- | --------------------------------- | -------------------------------------- | ----------------------------------------------------------- |
| Flat + floor | `Flat 4B, 2nd Floor` | `Flat 4 B, 2nd Floor` | Flat four B, second floor |
| House/door number | `No. 12/3, MG Road` | `Number 12 slash 3, M G Road` | Number twelve slash three, M G Road |
| Postal PIN in address | `Bengaluru - 560001` | `Bengaluru, 5 6 0 0 0 1` | Bengaluru, five six zero zero zero one |
| Full address | `42, MG Road, Bengaluru - 560001` | `42, M G Road, Bengaluru, 5 6 0 0 0 1` | forty two, M G Road, Bengaluru, five six zero zero zero one |
| Apartment unit | `B-204, Prestige Apts` | `B 204, Prestige Apts` | B two zero four, Prestige Apartments |
***
## 5. Aadhaar Card Numbers
Aadhaar is a 12-digit unique ID. Every digit must be separated by a single space. Strip any existing grouping spaces (4-4-4 format) before re-spacing individually.
| Scenario | Raw Input | Correct TTS Input | Expected Speech |
| -------------- | ---------------- | ------------------------- | ----------------------------------------------------------- |
| Full Aadhaar | `234567891234` | `2 3 4 5 6 7 8 9 1 2 3 4` | two three four five six seven eight nine one two three four |
| Grouped format | `2345 6789 1234` | `2 3 4 5 6 7 8 9 1 2 3 4` | *(strip groups first, same result)* |
| Masked Aadhaar | `XXXX XXXX 1234` | `X X X X X X X X 1 2 3 4` | X X X X X X X X one two three four |
***
## 6. Vehicle Registration Numbers
Indian vehicle registration numbers follow the format: **State Code** (2 letters) + **District Code** (2 digits) + **Series** (1–3 letters) + **Number** (4 digits). Each character is separated by a single space. Always strip hyphens before formatting. Letters must be uppercase.
**Format:** `[State Code] [District No.] [Series] [Reg. Number]`\
**Example:** `KA | 09 | MF | 1234` → `K A 0 9 M F 1 2 3 4`
| Scenario | Raw Input | Correct TTS Input | Expected Speech |
| ----------------- | --------------- | --------------------- | ------------------------------------ |
| Standard | `KA09MF1234` | `K A 0 9 M F 1 2 3 4` | K A zero nine M F one two three four |
| With hyphens | `KA-09-MF-1234` | `K A 0 9 M F 1 2 3 4` | K A zero nine M F one two three four |
| Delhi vehicle | `DL4CAF0001` | `D L 4 C A F 0 0 0 1` | D L four C A F zero zero zero one |
| Maharashtra | `MH12DE1433` | `M H 1 2 D E 1 4 3 3` | M H one two D E one four three three |
| Tamil Nadu | `TN22BC5678` | `T N 2 2 B C 5 6 7 8` | T N two two B C five six seven eight |
| Bharat series | `BH01AB1234` | `B H 0 1 A B 1 2 3 4` | B H zero one A B one two three four |
| 2-wheeler (short) | `KA05V1234` | `K A 0 5 V 1 2 3 4` | K A zero five V one two three four |
Do **not** pass `KA-09-MF-1234` with dashes and no spaces — the engine may misread the state code.
***
## 7. GST Numbers (GSTIN)
GSTIN is a 15-character alphanumeric code. Every character must be separated by a single space. Lowercase letters must be converted to uppercase. The letter `Z` at position 14 is always read as the letter **"Z"**, not **"zero"**.
**Structure:** `[State Code 2 digits]` + `[PAN 10 chars]` + `[Entity No.]` + `[Z]` + `[Check digit]`
| Scenario | Raw Input | Correct TTS Input | Expected Speech |
| ----------- | ----------------- | ------------------------------- | ---------------------------------------------------- |
| Karnataka | `29ABCDE1234F1Z5` | `2 9 A B C D E 1 2 3 4 F 1 Z 5` | two nine A B C D E one two three four F one Z five |
| Maharashtra | `27AABCU9603R1Z6` | `2 7 A A B C U 9 6 0 3 R 1 Z 6` | two seven A A B C U nine six zero three R one Z six |
| Delhi | `07AAACR5055K1ZR` | `0 7 A A A C R 5 0 5 5 K 1 Z R` | zero seven A A A C R five zero five five K one Z R |
| Tamil Nadu | `33AAACP1234C1Z1` | `3 3 A A A C P 1 2 3 4 C 1 Z 1` | three three A A A C P one two three four C one Z one |
Do **not** pass raw (`29ABCDE1234F1Z5`) or hyphenated (`29-ABCDE-1234-F1Z5`) formats — always space each character individually.
***
## 8. Dates
Accepted formats: `DD/MM/YY`, `DD-MM-YY`, `DD/MM/YYYY`, `DD-MM-YYYY`, `DD Month YYYY`, `DDth Month YYYY`. Do not mix separators. Use 4-digit years for historical dates.
| Format | Correct TTS Input | Expected Speech | Notes |
| ---------------- | ------------------ | ----------------------------------------- | ----------------------------------------------------- |
| `DD/MM/YY` | `15/08/47` | fifteenth August forty-seven | YY defaults to 2000s unless context implies otherwise |
| `DD-MM-YY` | `01-01-24` | first January twenty-four | |
| `DD/MM/YYYY` | `15/08/2024` | fifteenth August two thousand twenty-four | Preferred for unambiguous reading |
| `DD-MM-YYYY` | `01-01-2025` | first January two thousand twenty-five | |
| Natural language | `15th August 2024` | fifteenth August two thousand twenty-four | Use ordinal suffix always |
| Short month | `15 Aug 2024` | fifteenth August two thousand twenty-four | |
Do **not** mix separators — `15/08-2024` causes a parsing error and engine failure.
***
## 9. Currency and Prices
Indian Rupee must always be written as `Rs` — exact case, no period, no lowercase, no `INR`. Place the symbol before the amount with a single space. Use Indian comma grouping (`1,00,000`) for Rupee amounts — the engine reads lakh and crore natively from this format.
| Currency | Do NOT Use | Correct TTS Input | Expected Speech |
| -------------------- | -------------------------------------------- | ----------------- | ----------------------------- |
| Indian Rupee | `rs 500` / `RS. 500` / `Rs. 500` / `INR 500` | `Rs 500` | five hundred rupees |
| Rupee + paise | `Rs.49.50` / `rs 49.50` | `Rs 49.50` | forty-nine rupees fifty paise |
| Large amount (lakh) | `Rs 100000` | `Rs 1,00,000` | one lakh rupees |
| Large amount (crore) | `Rs 10000000` | `Rs 1,00,00,000` | one crore rupees |
| US Dollar | `dollar 20` / `USD20` | `$ 20` | twenty dollars |
| Euro | `euro 15` / `EUR15` | `€ 15` | fifteen euros |
| British Pound | `pound 30` / `GBP 30` | `£ 30` | thirty pounds |
| Japanese Yen | `yen 1000` / `JPY1000` | `¥ 1000` | one thousand yen |
***
## 10. Abbreviations and Acronyms
Acronyms and abbreviations must always be **UPPERCASE**. Lowercase or mixed-case versions risk being misread as words. Each uppercase letter is read individually.
| Type | Do NOT Use | Correct TTS Input | Expected Speech |
| -------------- | --------------- | ----------------- | --------------- |
| Organisation | `rnr` / `Rnr` | `RNR` | R N R |
| Paramilitary | `crpf` / `Crpf` | `CRPF` | C R P F |
| Space agency | `isro` / `Isro` | `ISRO` | I S R O |
| Finance term | `emi` / `Emi` | `EMI` | E M I |
| National ID | `pan` | `PAN` | P A N |
| Identity check | `kyc` / `Kyc` | `KYC` | K Y C |
| OTP/security | `otp` | `OTP` | O T P |
| Vehicle type | `suv` | `SUV` | S U V |
| Tax | `gst` | `GST` | G S T |
Some acronyms pronounced as words (NASA, NEFT, SWIFT) may be read as words rather than letter-by-letter. Verify pronunciation during QA for each acronym used in your content.
***
## 11. Ordinal Numbers
Always include the ordinal suffix (`st`, `nd`, `rd`, `th`). Without it, the engine reads bare numbers as cardinals ("one", "two") instead of ordinals ("first", "second").
| Context | Do NOT Use | Correct TTS Input | Expected Speech |
| ---------------- | -------------------- | ---------------------- | ------------------------ |
| Floor | `3 floor` | `3rd floor` | third floor |
| Rank | `1 rank` | `1st rank` | first rank |
| Date in sentence | `The 15 of August` | `The 15th of August` | the fifteenth of August |
| Anniversary | `25 anniversary` | `25th anniversary` | twenty-fifth anniversary |
| Queue position | `You are 2 in queue` | `You are 2nd in queue` | you are second in queue |
***
## 12. Decimals and Percentages
Use a period (`.`) as the decimal separator — never a comma. Percentages use the `%` symbol directly after the number with no space. Avoid trailing zeros in Rupee amounts unless paise must be communicated.
| Type | Do NOT Use | Correct TTS Input | Expected Speech |
| ------------------------- | ---------------- | ----------------- | ------------------------------------------- |
| Interest rate | `8,5%` | `8.5%` | eight point five percent |
| Discount | `20 percent off` | `20% off` | twenty percent off |
| Rupee (no trailing zeros) | `Rs 1,234.00` | `Rs 1,234` | one thousand two hundred thirty-four rupees |
| Weight/measure | `1,5 kg` | `1.5 kg` | one point five kilograms |
| Score/ratio | `7,8 out of 10` | `7.8 out of 10` | seven point eight out of ten |
***
## Quick Reference
| Data Type | Rule | Example |
| -------------------- | ----------------------------------------------------------------- | ---------------------------------- |
| Phone Number | Each digit separated by single space | `9 8 7 6 5 4 3 2 1 0` |
| Account Number | Each character (alpha + digit) separated by single space | `A B 1 2 3 4 5 6 7 8 9` |
| PIN Number | Each digit by single space; `0` = "zero" never "oh" | `4 8 2 1` |
| Aadhaar Number | Each of 12 digits separated by single space | `2 3 4 5 6 7 8 9 1 2 3 4` |
| Vehicle Reg. Number | Each character separated by single space; strip hyphens | `K A 0 9 M F 1 2 3 4` |
| GST / GSTIN | All 15 characters separated by single space; uppercase only | `2 9 A B C D E 1 2 3 4 F 1 Z 5` |
| Address (postal PIN) | PIN digits only → digit-by-digit; house numbers read naturally | `Bengaluru, 5 6 0 0 0 1` |
| Currency (Rupee) | Use `Rs` — not `rs` / `RS` / `Rs.` / `INR`; Indian comma grouping | `Rs 1,00,000` |
| Currency (others) | Use symbol `$` `€` `£` `¥` before amount with one space | `$ 20` / `€ 15` |
| Dates | `DD/MM/YYYY` or `DD-MM-YYYY`; ordinal suffix in natural language | `15/08/2024` or `15th August 2024` |
| Acronyms | Always UPPERCASE — never lowercase or mixed case | `CRPF` `KYC` `OTP` |
| Ordinals | Always include suffix: `1st`, `2nd`, `3rd`, `15th` | `3rd floor` / `15th August` |
| Decimals | Period (`.`) not comma; `%` directly after number | `8.5%` / `1.5 kg` |
# Text-to-Speech (Streaming)
Source: https://docs.gnani.ai/api/TTS/tts-sse
POST /api/v1/tts/sse
Stream audio in chunks as it's generated via Server-Sent Events.
**Currently in beta.** You're on the priority waitlist and among the first to get access.
## Overview
Receive audio in chunks as it's generated, allowing playback to start immediately. Reduces latency compared to [TTS REST](/vachana/TTS/tts-inference).
Passing numbers, IDs, dates, or currency as raw strings causes mispronunciations. See the [Input Formatting Guide](/vachana/TTS/tts-input-formating) for correct formatting of phone numbers, account numbers, PINs, Aadhaar, vehicle registration numbers, GSTIN, currency, and more.
***
## Available Voices
| Voice | Gender | Description |
| :------ | :----- | :----------------------- |
| Pranav | Male | Bold, Trustworthy |
| Kaveri | Female | Confident, Bright |
| Shubhra | Female | Gentle, Expressive |
| Deepak | Male | Grounded, Conversational |
***
## Python SDK
The SDK's streaming client handles SSE parsing and chunk reassembly for you — you just iterate and write.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
```python Constructor argument theme={null}
from gnani.tts import GnaniTTSStreamClient
client = GnaniTTSStreamClient(api_key="your-api-key")
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.tts import GnaniTTSStreamClient
client = GnaniTTSStreamClient()
```
### Stream Audio to a File
`synthesize_stream` yields audio chunks as they arrive. Playback or writing can begin before the full response is complete.
```python theme={null}
from gnani.tts import GnaniTTSStreamClient
client = GnaniTTSStreamClient(api_key="your-api-key")
with open("output.wav", "wb") as f:
for chunk in client.synthesize_stream(
"Streaming TTS response in Hindi",
voice="sia",
):
f.write(chunk)
```
### With Custom Audio Config
```python theme={null}
from gnani.tts import GnaniTTSStreamClient, AudioConfig
client = GnaniTTSStreamClient(api_key="your-api-key")
with open("output.wav", "wb") as f:
for chunk in client.synthesize_stream(
"नमस्ते, आप कैसे हैं?",
voice="raju",
audio_config=AudioConfig(
sample_rate=44100,
encoding="linear_pcm",
container="wav",
),
):
f.write(chunk)
```
## Supported Languages
The Gnani Timbre v2.0 API supports 2 languages.
| Language | Native Script | Example |
| -------- | ------------------- | -------------------------- |
| English | Latin | "I am going to the market" |
| Hindi | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
# Text-to-Speech (Realtime)
Source: https://docs.gnani.ai/api/TTS/tts-websocket
Real-time text-to-speech with streaming audio via WebSocket.
**Currently in beta.** You're on the priority waitlist and among the first to get access.
## Overview
Stream audio in real-time with the lowest latency. Perfect for interactive assistants and live applications. For simpler use cases, see [TTS REST](/vachana/TTS/tts-inference) or [TTS SSE](/vachana/TTS/tts-sse).
Passing numbers, IDs, dates, or currency as raw strings causes mispronunciations. See the [Input Formatting Guide](/vachana/TTS/tts-input-formating) for correct formatting of phone numbers, account numbers, PINs, Aadhaar, vehicle registration numbers, GSTIN, currency, and more.
## Available Voices
| Voice | Gender | Description |
| :------ | :----- | :----------------------- |
| Pranav | Male | Bold, Trustworthy |
| Kaveri | Female | Confident, Bright |
| Shubhra | Female | Gentle, Expressive |
| Deepak | Male | Grounded, Conversational |
## Endpoint
```text theme={null}
wss://api.vachana.ai/api/v1/tts
```
## Authentication
All Realtime connections require the following headers:
| Header | Required | Description | Example |
| -------------- | -------- | ------------------------------- | ------------------- |
| `Content-Type` | Yes | Must be `application/json` | `application/json` |
| `X-API-Key-ID` | Yes | Your API key for authentication | `` |
## Request Format
Send a JSON message with the following structure:
```json theme={null}
{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-voice-v3",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm"
}
}
```
Number of audio channels (e.g., `1` for mono, `2` for stereo)
Sample width in bytes (e.g., `2` for 16-bit audio)
Audio encoding format (e.g., `linear_pcm`)
Audio container format (e.g., `wav`)
## Response
The server streams audio data in real-time as binary chunks. Each chunk contains PCM audio data according to the specified `audio_config`.
## Example Usage
```javascript JavaScript theme={null}
const ws = new WebSocket("wss://api.vachana.ai/api/v1/tts", {
headers: {
"Content-Type": "application/json",
"X-API-Key-ID": "",
},
});
ws.on("open", () => {
const request = {
text: "नमस्ते, आप कैसे हैं?",
model: "vachana-voice-v3",
audio_config: {
sample_rate: 44100,
encoding: "linear_pcm",
},
};
ws.send(JSON.stringify(request));
});
ws.on("message", (data) => {
// Handle audio chunks
console.log("Received audio chunk:", data);
});
ws.on("error", (error) => {
console.error("WebSocket error:", error);
});
ws.on("close", () => {
console.log("WebSocket connection closed");
});
```
```python Python theme={null}
import websocket
import json
def on_message(ws, message):
# Handle audio chunks
print(f"Received audio chunk: {len(message)} bytes")
def on_error(ws, error):
print(f"Error: {error}")
def on_close(ws, close_status_code, close_msg):
print("WebSocket connection closed")
def on_open(ws):
request = {
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-voice-v3",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm"
}
}
ws.send(json.dumps(request))
ws = websocket.WebSocketApp(
"wss://api.vachana.ai/api/v1/tts",
header={
"Content-Type": "application/json",
"X-API-Key-ID": ""
},
on_open=on_open,
on_message=on_message,
on_error=on_error,
on_close=on_close
)
ws.run_forever()
```
***
## Python SDK
The SDK's realtime client manages the WebSocket lifecycle, audio streaming, and async iteration so you can focus on your application logic.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
```python Constructor argument theme={null}
from gnani.tts import GnaniTTSRealtimeClient
client = GnaniTTSRealtimeClient(api_key="your-api-key")
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.tts import GnaniTTSRealtimeClient
client = GnaniTTSRealtimeClient()
```
### Stream Audio Chunks in Real-Time
Use the async context manager to open the connection and iterate over audio chunks as they arrive.
```python theme={null}
import asyncio
from gnani.tts import GnaniTTSRealtimeClient
async def main():
async with GnaniTTSRealtimeClient(api_key="your-api-key") as client:
with open("output.wav", "wb") as f:
async for chunk in client.synthesize(
"नमस्ते, आप कैसे हैं?",
voice="sia",
):
f.write(chunk)
asyncio.run(main())
```
### Collect All Audio at Once
If you don't need to process chunks as they arrive, use `synthesize_and_collect` to get the full audio as a single bytes object.
```python theme={null}
import asyncio
from gnani.tts import GnaniTTSRealtimeClient
async def main():
async with GnaniTTSRealtimeClient(api_key="your-api-key") as client:
audio = await client.synthesize_and_collect(
"Realtime TTS response",
voice="neha",
)
with open("output.wav", "wb") as f:
f.write(audio)
asyncio.run(main())
```
## Supported Languages
The Gnani Timbre v2.0 API supports 2 languages.
| Language | Native Script | Example |
| -------- | ------------------- | -------------------------- |
| English | Latin | "I am going to the market" |
| Hindi | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
# Voice Cloned TTS (REST)
Source: https://docs.gnani.ai/api/VC/vc-inference
POST /api/v1/tts/inference
Generate cloned voice audio in a single synchronous response.
## Overview
Upload a reference audio clip to get a `speaker_embedding`. See [Voice Clone Embeddings](/vachana/VC/voice-clone-embeddings).
Pass the `speaker_embedding` from Step 1 to this endpoint to synthesize audio in your cloned voice.
Synthesize audio using your cloned voice. Pass the `speaker_embedding` obtained from [Voice Clone Embeddings](/vachana/VC/voice-clone-embeddings) along with your text. The full audio is returned in one response. For streaming playback, see [Voice Cloning Streaming](/vachana/VC/vc-sse) or [Voice Cloning Realtime](/vachana/VC/vc-websocket).
## Endpoint
```
POST https://api.vachana.ai/api/v1/tts/inference
```
## Authentication
| Header | Required | Description | Example |
| -------------- | -------- | ------------------------------- | ------------------ |
| `X-API-Key-ID` | Yes | Your API key for authentication | `your-api-key-id` |
| `Content-Type` | Yes | Must be `application/json` | `application/json` |
## Request Body
The text to synthesize into speech
Voice cloning model to use. Currently supported: `vachana-vc-v1`
Audio output configuration
Sample rate in Hz (8000-44100)
Number of audio channels (1-8)
Sample width in bytes (1-4)
Audio encoding format: `linear_pcm` or `oggopus`
Audio container format: `raw`, `mp3`, `wav`, `mulaw`, or `ogg`
MP3 bitrate (only when container=mp3): `96k`, `128k`, or `192k`
Voice clone embedding obtained from the Voice Clone Embeddings endpoint
The voice clone embedding string
Shape of the embedding tensor, e.g., `[1, 768]`
Data type of the embedding, e.g., `torch.bfloat16`
## Response
Returns binary audio data in the format specified by `audio_config.container`:
* `audio/wav` for WAV files
* `audio/mpeg` for MP3 files
* `audio/ogg` for OGG files
## Example Request
```bash cURL theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/inference \
-H "X-API-Key-ID: your-api-key-id" \
-H "Content-Type: application/json" \
-d '{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm",
"container": "wav"
},
"speaker_embedding": {
"embedding": "your-embedding-string",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}' \
--output audio.wav
```
```python Python theme={null}
import requests
url = "https://api.vachana.ai/api/v1/tts/inference"
headers = {
"X-API-Key-ID": "your-api-key-id",
"Content-Type": "application/json"
}
payload = {
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm",
"container": "wav"
},
"speaker_embedding": {
"embedding": "your-embedding-string",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}
response = requests.post(url, headers=headers, json=payload)
if response.status_code == 200:
with open("audio.wav", "wb") as f:
f.write(response.content)
print("Audio saved successfully")
else:
print(f"Error: {response.status_code}")
print(response.json())
```
```javascript JavaScript theme={null}
const url = "https://api.vachana.ai/api/v1/tts/inference";
const headers = {
"X-API-Key-ID": "your-api-key-id",
"Content-Type": "application/json"
};
const payload = {
text: "नमस्ते, आप कैसे हैं?",
model: "vachana-vc-v1",
audio_config: {
sample_rate: 44100,
encoding: "linear_pcm",
container: "wav"
},
speaker_embedding: {
embedding: "your-embedding-string",
shape: [1, 768],
dtype: "torch.bfloat16"
}
};
fetch(url, {
method: "POST",
headers: headers,
body: JSON.stringify(payload)
})
.then(response => response.blob())
.then(blob => {
const url = window.URL.createObjectURL(blob);
const a = document.createElement("a");
a.href = url;
a.download = "audio.wav";
a.click();
})
.catch(error => console.error("Error:", error));
```
## Error Responses
Invalid text or audio configuration
```json theme={null}
{
"success": false,
"error": {
"type": "INVALID_REQUEST_ERROR",
"message": "Invalid text or audio configuration."
}
}
```
Rate limit exceeded
```json theme={null}
{
"success": false,
"error": {
"type": "RATE_LIMIT_ERROR",
"message": "Rate limit exceeded. Please try again later."
}
}
```
Unexpected error occurred
```json theme={null}
{
"success": false,
"error": {
"type": "API_ERROR",
"message": "An unexpected error occurred while processing."
}
}
```
# Voice Cloned TTS (Streaming)
Source: https://docs.gnani.ai/api/VC/vc-sse
POST /api/v1/tts/sse
Stream cloned voice audio in chunks via Server-Sent Events.
## Overview
Upload a reference audio clip to get a `speaker_embedding`. See [Voice Clone Embeddings](/vachana/VC/voice-clone-embeddings).
Pass the `speaker_embedding` from Step 1 to this endpoint to stream cloned voice audio progressively.
Stream cloned voice audio as it's generated using Server-Sent Events. Pass the `speaker_embedding` from [Voice Clone Embeddings](/vachana/VC/voice-clone-embeddings) to use your cloned voice. Reduces latency compared to [Voice Cloned TTS REST](/vachana/VC/vc-inference). For the lowest latency, see [Voice Cloned TTS Realtime](/vachana/VC/vc-websocket).
## Endpoint
```
POST https://api.vachana.ai/api/v1/tts/sse
```
## Authentication
| Header | Required | Description | Example |
| -------------- | -------- | ------------------------------- | ------------------ |
| `X-API-Key-ID` | Yes | Your API key for authentication | `your-api-key-id` |
| `Content-Type` | Yes | Must be `application/json` | `application/json` |
## Request Body
The text to synthesize into speech
Voice cloning model to use. Currently supported: `vachana-vc-v1`
Audio output configuration
Sample rate in Hz (8000-44100)
Number of audio channels (1-8)
Sample width in bytes (1-4)
Audio encoding format: `linear_pcm` or `oggopus`
Audio container format: `raw`, `mp3`, `wav`, `mulaw`, or `ogg`
MP3 bitrate (only when container=mp3): `96k`, `128k`, or `192k`
Voice clone embedding obtained from the Voice Clone Embeddings endpoint
The voice clone embedding string
Shape of the embedding tensor, e.g., `[1, 768]`
Data type of the embedding, e.g., `torch.bfloat16`
## Response
The server streams audio data via Server-Sent Events (SSE). Each event contains a chunk of audio data encoded in base64.
### Event Types
Contains base64-encoded audio data
```
event: audio_chunk
data: UklGRiQAAABXQVZFZm10IBAAAAABAAEAQB8AAEAfAAABAAgAZGF0YQAAAAA=
```
Signals the end of the audio stream
```
event: completed
data: {"status": "success"}
```
## Example Request
```bash cURL theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/sse \
-H "X-API-Key-ID: your-api-key-id" \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-N \
-d '{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm",
"container": "wav"
},
"speaker_embedding": {
"embedding": "your-embedding-string",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}'
```
```python Python theme={null}
import requests
import base64
url = "https://api.vachana.ai/api/v1/tts/sse"
headers = {
"X-API-Key-ID": "your-api-key-id",
"Content-Type": "application/json",
"Accept": "text/event-stream"
}
payload = {
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm",
"container": "wav"
},
"speaker_embedding": {
"embedding": "your-embedding-string",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}
response = requests.post(url, headers=headers, json=payload, stream=True)
audio_chunks = []
for line in response.iter_lines():
if line:
line = line.decode('utf-8')
if line.startswith('data: '):
data = line[6:]
if data.startswith('{'):
# Completed event
print("Stream completed")
else:
# Audio chunk
audio_chunks.append(base64.b64decode(data))
# Combine and save audio
with open("audio.wav", "wb") as f:
for chunk in audio_chunks:
f.write(chunk)
print("Audio saved successfully")
```
```javascript JavaScript theme={null}
const url = "https://api.vachana.ai/api/v1/tts/sse";
const headers = {
"X-API-Key-ID": "your-api-key-id",
"Content-Type": "application/json",
"Accept": "text/event-stream"
};
const payload = {
text: "नमस्ते, आप कैसे हैं?",
model: "vachana-vc-v1",
audio_config: {
sample_rate: 44100,
encoding: "linear_pcm",
container: "wav"
},
speaker_embedding: {
embedding: "your-embedding-string",
shape: [1, 768],
dtype: "torch.bfloat16"
}
};
const eventSource = new EventSource(url);
const audioChunks = [];
eventSource.addEventListener("audio_chunk", (event) => {
const audioData = atob(event.data);
const bytes = new Uint8Array(audioData.length);
for (let i = 0; i < audioData.length; i++) {
bytes[i] = audioData.charCodeAt(i);
}
audioChunks.push(bytes);
});
eventSource.addEventListener("completed", (event) => {
console.log("Stream completed");
// Combine chunks and create blob
const blob = new Blob(audioChunks, { type: "audio/wav" });
const url = window.URL.createObjectURL(blob);
const a = document.createElement("a");
a.href = url;
a.download = "audio.wav";
a.click();
eventSource.close();
});
eventSource.onerror = (error) => {
console.error("SSE Error:", error);
eventSource.close();
};
// Send the request
fetch(url, {
method: "POST",
headers: headers,
body: JSON.stringify(payload)
});
```
## Error Responses
Invalid text or audio configuration
```json theme={null}
{
"success": false,
"error": {
"type": "INVALID_REQUEST_ERROR",
"message": "Invalid text or audio configuration."
}
}
```
Rate limit exceeded
```json theme={null}
{
"success": false,
"error": {
"type": "RATE_LIMIT_ERROR",
"message": "Rate limit exceeded. Please try again later."
}
}
```
Unexpected error occurred
```json theme={null}
{
"success": false,
"error": {
"type": "API_ERROR",
"message": "An unexpected error occurred while processing."
}
}
```
# Voice Cloned TTS (Realtime)
Source: https://docs.gnani.ai/api/VC/vc-websocket
Real-time cloned voice audio streaming via WebSocket.
## Overview
Upload a reference audio clip to get a `speaker_embedding`. See [Voice Clone Embeddings](/vachana/VC/voice-clone-embeddings).
Pass the `speaker_embedding` from Step 1 to this endpoint to stream cloned voice audio in real-time with the lowest latency.
Stream cloned voice audio in real-time with the lowest latency via WebSocket. Pass the `speaker_embedding` from [Voice Clone Embeddings](/vachana/VC/voice-clone-embeddings) to use your cloned voice. For simpler use cases, see [Voice Cloning REST](/vachana/VC/vc-inference) or [Voice Cloning Streaming](/vachana/VC/vc-sse).
## Endpoint
```
wss://api.vachana.ai/api/v1/tts
```
## Authentication
All Realtime connections require the following headers:
| Header | Required | Description | Example |
| -------------- | -------- | ------------------------------- | ------------------- |
| `Content-Type` | Yes | Must be `application/json` | `application/json` |
| `X-API-Key-ID` | Yes | Your API key for authentication | `` |
## Request Format
Send a JSON message with the following structure:
```json theme={null}
{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm"
},
"speaker_embedding": {
"embedding": "",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}
```
## Response
The server streams audio data in real-time as binary chunks. Each chunk contains PCM audio data according to the specified `audio_config`.
## Example Usage
```javascript JavaScript theme={null}
const ws = new WebSocket("wss://api.vachana.ai/api/v1/tts", {
headers: {
"Content-Type": "application/json",
"X-API-Key-ID": "",
},
});
ws.on("open", () => {
const request = {
text: "नमस्ते, आप कैसे हैं?",
model: "vachana-vc-v1",
audio_config: {
sample_rate: 44100,
encoding: "linear_pcm",
},
speaker_embedding: {
embedding: "",
shape: [1, 768],
dtype: "torch.bfloat16",
},
};
ws.send(JSON.stringify(request));
});
ws.on("message", (data) => {
// Handle audio chunks
console.log("Received audio chunk:", data);
});
ws.on("error", (error) => {
console.error("WebSocket error:", error);
});
ws.on("close", () => {
console.log("WebSocket connection closed");
});
```
```python Python theme={null}
import websocket
import json
def on_message(ws, message):
# Handle audio chunks
print(f"Received audio chunk: {len(message)} bytes")
def on_error(ws, error):
print(f"Error: {error}")
def on_close(ws, close_status_code, close_msg):
print("WebSocket connection closed")
def on_open(ws):
request = {
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm"
},
"speaker_embedding": {
"embedding": "",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}
ws.send(json.dumps(request))
ws = websocket.WebSocketApp(
"wss://api.vachana.ai/api/v1/tts",
header={
"Content-Type": "application/json",
"X-API-Key-ID": ""
},
on_open=on_open,
on_message=on_message,
on_error=on_error,
on_close=on_close
)
ws.run_forever()
```
# Voice Clone Embeddings
Source: https://docs.gnani.ai/api/VC/voice-clone-embeddings
POST /api/v1/tts/voice-clone/embeddings
Generate voice clone embeddings from an audio file.
## Voice Cloning Flow
Voice cloning is a two-step process. Complete Step 1 once per voice, then reuse the embedding across any synthesis endpoint.
Upload 5–30 seconds of clean reference audio to extract a `speaker_embedding`. Cache the result — you only need to generate it once per voice.
Pass the `speaker_embedding` from Step 1 to your preferred synthesis endpoint:
Full audio returned in a single response
Receive audio progressively as it's synthesized
Lowest latency — stream text in, audio out
***
## Overview
Generate a `speaker_embedding` from a reference audio clip. Upload the file and receive a multi-dimensional embedding you can pass to any Voice Cloned TTS endpoint.
# Introduction
Source: https://docs.gnani.ai/api/introduction/introduction
[Gnani AI](https://app.gnani.ai/voice) provides high-accuracy speech-to-text and text-to-speech APIs across 10+ Indian languages. Every API is optimized for Indian speech patterns, including regional dialects, code-switching, and accurate native-script transcription.
## Available APIs
### Speech-to-Text: Gnani Prisma v2.5
| API | Description |
| ---------------- | ------------------------------------------------------------------------------------------ |
| **STT REST** | Transcribe short audio files (≤ 60s) via a single HTTP request |
| **STT Realtime** | Stream live audio over a WebSocket connection and receive transcript segments in real-time |
| **STT Batch** | Submit long or multiple audio files for async transcription; poll for results via `job_id` |
### Text-to-Speech: Gnani Timbre v2.0
| API | Description |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------- |
| **TTS REST** | Synthesize text to audio in a single synchronous HTTP call |
| **TTS Streaming** | Submit text via an HTTP request and receive synthesized audio progressively as a server-sent event stream |
| **TTS Realtime** | Stream text incrementally and receive audio simultaneously over a persistent WebSocket connection, delivering low latency |
***
### Voice Cloning
| API | Description |
| ------------------------------ | ---------------------------------------------------------------------------------------- |
| **VC Embeddings** | Upload a reference audio file to generate a `speaker_embedding` for use in voice cloning |
| **Voice Cloned TTS REST** | Synthesize audio in your cloned voice via a single synchronous HTTP call |
| **Voice Cloned TTS Streaming** | Stream cloned voice audio progressively using Server-Sent Events |
| **Voice Cloned TTS Realtime** | Stream text and receive cloned voice audio in real-time over a WebSocket connection |
## Key Capabilities
| Feature | Detail |
| ------------------------- | ---------------------------------------------------------------------------------- |
| **10+ Indian Languages** | Native script transcription and synthesis across 10+ Indian languages |
| **Language Detection** | Automatic — or specify `language_code` to target a specific language. |
| **Code-Switching** | Handles code-mixed speech naturally |
| **Audio Flexibility** | STT accepts WAV, MP3, OGG, FLAC, AAC, M4A |
| **Voice Cloning** | Clone any voice from a short audio sample using speaker embeddings |
| **Latency** | P95 200ms for STT
Streaming TTS supported |
| **Processing Modes** | Real-time streaming and batch |
| **Accuracy** | Sub-4% WER on Indian English and 20-30% better accuracy for major Indian languages |
| **Transcript Formatting** | Auto-punctuation, inverse text normalization (numerals, dates, currency) |
| **SSML Support** | Full SSML for fine-grained speech synthesis control |
| **Noise Robustness** | Optimized for telephony-grade and noisy real-world audio |
***
## Get Started
Ready to begin? Head over to the [Quick Start Guide](/vachana/introduction/quick-start) to make your first API call.
# Quick Start
Source: https://docs.gnani.ai/api/introduction/quick-start
Make your first speech-to-text, text-to-speech or voice-cloned TTS API request in minutes.
## Prerequisites
Before you begin, ensure you have:
* A valid API key (sign up on the [Gnani API platform](https://app.gnani.ai/voice/) to generate API keys)
* cURL installed, or an API client such as Postman
Use a test audio file with the following requirements -
1. Format: WAV, MP3, OGG, FLAC, AAC, M4A
2. Sampling rate: 8 kHz – 44.1 kHz
3. Maximum duration: 60 seconds
### Your First Speech-to-Text Request
Minimal example to transcribe a Hindi audio file:
```bash theme={null}
curl -X POST https://api.vachana.ai/stt/v3 \
-H 'Content-Type: multipart/form-data' \
-H 'X-API-Key-ID: ' \
-F audio_file='/path/to/your/audio.wav' \
-F language_code=hi-IN
```
**Replace these values:**
* ``: Your Gnani Prisma v2.5 API key
* `/path/to/your/audio.wav`: Path to your audio file
* `hi-IN`: Language code (see [Language Codes](/vachana/STT/speech-to-text#language-codes) for all options)
### Expected STT Response
On success, you'll receive a JSON response like:
```json theme={null}
{
"success": true,
"transcript": "नमस्ते, आप कैसे हैं?"
}
```
For real-time, low-latency transcription of streaming audio, Gnani Prisma v2.5 provides a Realtime API.
### Connection
Create a Realtime connection using your API credentials:
```javascript theme={null}
const ws = new Realtime("wss://api.vachana.ai/stt/v3", {
headers: {
"x-api-key-id": "",
}
});
```
### Send Audio
Send raw PCM audio frames over the Realtime connection. Audio requirements:
* Format: PCM 16-bit
* Sample rate: 16 or 8 kHz
* Channels: Mono
* Chunk size: **1024 bytes per frame** (512 samples)
Client-to-server messages must contain **only binary audio frames**. Do **not** wrap the audio in JSON.
### Expected STT Response
The server sends JSON text frames containing transcription segments:
```json theme={null}
{
"type": "transcript",
"timestamp": "2024-01-15T10:30:05.987Z",
"text": "Hello, how are you today?",
"audio_duration_ms": 2340,
"segment_id": "seg_abc123",
"segment_index": 1,
"latency": 320,
"detected_language": "en"
}
```
Have your input text ready. You'll also need a voice name — see Voice Options for available voices.
### Your First Text-to-Speech Call
Minimal example for REST TTS (synchronous audio). This endpoint returns the full synthesized audio as a binary response.
```bash theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/inference \
-H 'Content-Type: application/json' \
-H 'X-API-Key-ID: ' \
-d '{
"text": "नमस्ते, आप कैसे हैं?",
"voice": "sia",
"model": "vachana-voice-v2",
"audio_config": {
"sample_rate": 44100,
"num_channels": 1,
"sample_width": 2,
"encoding": "linear_pcm",
"container": "wav"
}
}' \
-output response.wav
```
### Expected TTS Response
A successful request will return a `200 OK` HTTP status. The response body will contain raw binary audio data representing the synthesized text, adhering to the format specified in your `audio_config`.
```http theme={null}
HTTP/1.1 200 OK
Content-Type: audio/wav
```
This endpoint streams synthesized audio using Server-Sent Events (SSE). Audio is generated and delivered incrementally as it becomes available.
### Your First Streaming Call
```bash theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/sse \
-H 'Content-Type: application/json' \
-H 'X-API-Key-ID: ' \
-d '{
"text": "नमस्ते, आप कैसे हैं?",
"voice": "sia",
"model": "vachana-voice-v2"
}'
```
### Expected SSE Response
A successful request will return a `200 OK` HTTP status. The response body will contain a stream of server-sent events. Each chunk contains base64 encoded audio fragments.
```http theme={null}
HTTP/1.1 200 OK
Content-Type: text/event-stream
event: audio_chunk
data: UklGRiQAAABXQVZFZm10IBAAAAABAAEAQB8AAEAfAAABAAgAZGF0YQAAAAA=
event: audio_chunk
data: //NkxAAAAANIAAAAAExBTUUzLjEwMKqqqqqqqqqqqqqqqqqqqqqqqqqqqqqq
event: completed
data: {"status": "success"}
```
For ultra-low latency applications, the Realtime API allows you to stream text input and receive synthesized audio continuously.
### Connection
Connect using a Realtime client with the required authentication headers:
```javascript theme={null}
const ws = new Realtime("wss://api.vachana.ai/api/v1/tts", {
headers: {
"x-api-key-id": "",
}
});
```
### Send Text
Once connected, send a JSON payload containing the text to synthesize.
```json theme={null}
{
"text": "नमस्ते, आप कैसे हैं?",
"voice": "sia",
"model": "vachana-voice-v2"
}
```
### Audio Stream Response
Upon successful connection, the server will return a `101 Switching Protocols` status to establish the Realtime.
Once the text payload is sent, the server will immediately begin streaming audio back as a sequence of **binary Realtime frames** containing the raw PCM audio data, terminating or keeping the connection open depending on the application context.
Voice cloning works in two steps:
1. **Generate embeddings** — upload a reference audio file to get a `speaker_embedding`
2. **Synthesize** — pass the embedding with your text to any VC TTS endpoint
### Step 1: Generate Voice Embeddings
Upload a reference audio file (WAV/MP3, ideally 5–30 seconds of clear speech):
```bash theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/voice-clone/embeddings \
-H 'X-API-Key-ID: ' \
-F audio_file='@/path/to/reference.wav'
```
**Replace these values:**
* ``: Your Gnani Timbre v2.0 API key
* `/path/to/reference.wav`: Path to your reference audio file
### Expected Embeddings Response
```json theme={null}
{
"embedding": "",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
```
### Step 2: Synthesize with Your Cloned Voice
Pass the `speaker_embedding` from Step 1 to synthesize audio in your cloned voice:
```bash theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/inference \
-H 'Content-Type: application/json' \
-H 'X-API-Key-ID: ' \
-d '{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"audio_config": {
"sample_rate": 44100,
"num_channels": 1,
"sample_width": 2,
"encoding": "linear_pcm",
"container": "wav"
},
"speaker_embedding": {
"embedding": "",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}' \
--output cloned_voice.wav
```
A successful request returns a `200 OK` with raw binary audio data in the specified format.
Stream cloned voice audio progressively via Server-Sent Events:
```bash theme={null}
curl -X POST https://api.vachana.ai/api/v1/tts/sse \
-H 'Content-Type: application/json' \
-H 'X-API-Key-ID: ' \
-d '{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-vc-v1",
"speaker_embedding": {
"embedding": "",
"shape": [1, 768],
"dtype": "torch.bfloat16"
}
}'
```
The response streams base64-encoded audio chunks as server-sent events, identical in format to the TTS SSE response.
For the lowest latency, stream text and receive cloned voice audio over a WebSocket:
```javascript theme={null}
const ws = new WebSocket("wss://api.vachana.ai/api/v1/tts", {
headers: {
"Content-Type": "application/json",
"X-API-Key-ID": "",
},
});
ws.on("open", () => {
ws.send(JSON.stringify({
text: "नमस्ते, आप कैसे हैं?",
model: "vachana-vc-v1",
audio_config: { sample_rate: 44100, encoding: "linear_pcm" },
speaker_embedding: {
embedding: "",
shape: [1, 768],
dtype: "torch.bfloat16",
},
}));
});
ws.on("message", (data) => {
// Handle binary PCM audio chunks
});
```
The server streams binary PCM audio chunks over the WebSocket connection.
***
## Next Steps
* **Speech-to-Text**: [STT REST](/vachana/STT/speech-to-text#api-reference) and [STT Realtime](/vachana/STT/stt-websocket) for all STT parameters and language options
* **Text-to-Speech**: [REST](/vachana/TTS/tts-inference), [Streaming (SSE)](/vachana/TTS/tts-sse), and [Realtime](/vachana/TTS/tts-websocket) for TTS options
* **Voice Cloning**: [VC Embeddings](/vachana/VC/voice-clone-embeddings), [REST](/vachana/VC/vc-inference), [Streaming](/vachana/VC/vc-sse), and [Realtime](/vachana/VC/vc-websocket) for voice cloning options
# Call Analytics Pipeline
Source: https://docs.gnani.ai/api/use-cases/call-analytics
A production-ready Python pipeline for call analytics using the Gnani Prisma v2.5 Batch STT API. Covers async transcription with two-speaker diarization, segment-level sentiment, and LLM-powered analysis via Claude or OpenAI.
## Overview
Every customer call contains signal that most teams never act on. This pipeline surfaces that signal automatically — agent effectiveness, customer sentiment, resolution quality, and upsell opportunities — across any volume of calls, in 10 Indian languages.
| Industry | What the pipeline enables |
| -------------------------- | ------------------------------------------------------------------------------------------------------ |
| **BFSI / Collections** | Monitor agent compliance, detect customer distress early, flag missed EMI restructuring opportunities. |
| **Insurance** | Analyze claim support calls, track resolution rates, identify policy renewal signals. |
| **Contact Centers / BPOs** | Automate QA at scale, reduce manual call review, improve agent training with structured feedback. |
| **Healthcare** | Analyze patient support calls, surface unresolved queries, track sentiment across touchpoints. |
| **Telecom** | Detect churn signals, identify upsell triggers, monitor service complaint patterns. |
**Native sentiment — no extra inference cost.** The Gnani Prisma v2.5 Batch STT API returns `sentiment` and `emotion` per segment directly from the transcription layer. You get speaker-wise sentiment timelines without relying solely on the LLM analysis step.
***
## Prerequisites & Installation
```bash theme={null}
# Gnani Vachana SDK
pip install gnani-vachana
# Audio chunking (required for files over 1 hour)
pip install pydub
# Install whichever LLM provider you intend to use
pip install anthropic # Claude
pip install openai # OpenAI / ChatGPT
```
**ffmpeg required for non-WAV formats.** pydub needs ffmpeg to process MP3, AAC, M4A, and OGG files. Install it with `brew install ffmpeg` on macOS or `apt install ffmpeg` on Linux.
***
## Authentication
Every request to the Batch STT API requires the `X-API-Key-ID` header. Store all credentials in environment variables.
| Header | Required | Description |
| ------------------ | -------- | ---------------------------------------------------------------------------------------------- |
| `X-API-Key-ID` | Yes | Your Gnani Prisma v2.5 API key. Required on every request — both submit and status calls. |
| `X-API-Request-ID` | No | A UUID trace ID you assign. Used to correlate your logs with platform logs or support tickets. |
```bash .env theme={null}
# Vachana
GNANI_API_KEY=your-api-key
# LLM provider — set whichever you will use
ANTHROPIC_API_KEY=your-anthropic-key
OPENAI_API_KEY=your-openai-key
# Switch between "claude" and "openai"
LLM_PROVIDER=claude
```
**Never hardcode API keys.** Do not commit API keys to version control. Use environment variables, a secrets manager, or a vault. Rotate keys immediately if exposed.
***
## Supported Languages
| Language | Code | Script | ITN Support |
| ------------- | ------------- | ------------------ | ------------ |
| **Hindi** | `hi-IN` | Devanagari | Yes |
| **English** | `en-IN` | Latin | Yes |
| **Tamil** | `ta-IN` | Tamil | — |
| **Telugu** | `te-IN` | Telugu | — |
| **Kannada** | `kn-IN` | Kannada | — |
| **Malayalam** | `ml-IN` | Malayalam | — |
| **Bengali** | `bn-IN` | Bengali | — |
| **Gujarati** | `gu-IN` | Gujarati | — |
| **Marathi** | `mr-IN` | Devanagari | — |
| **Punjabi** | `pa-IN` | Gurmukhi | — |
| **Hinglish** | `en-hi-in-cm` | Latin + Devanagari | Experimental |
ITN converts spoken-form numbers, currency, dates, and phone numbers into written form — for example, "five thousand rupees" becomes "₹5,000". Set `format=transcribe` in the request to enable it.
***
## Batch API Flow
| Operation | Method | Endpoint |
| ---------------- | ------ | ----------------------------------------------------- |
| **Submit job** | `POST` | `https://api.vachana.ai/stt/v3/batch/submit` |
| **Check status** | `GET` | `https://api.vachana.ai/stt/v3/batch/status/{job_id}` |
Upload 1–10 audio files as multipart form data with your language code and format preference. Receive a `job_id` immediately. Transcription has not started at this point.
Persist the `job_id` from the submit response. It is required for every subsequent status call.
Call the status endpoint every 60 seconds. Status transitions: `submitted` → `processing` → `completed` or `failed`. The `results` field is `null` until the job reaches `completed`.
Extract per-segment fields: `speaker_id`, `text`, `start_time`, `end_time`, `sentiment`, `emotion`, `confidence`. Build speaker-separated conversation threads and talk-time logs.
Send the parsed transcript to Claude or OpenAI with a structured analysis prompt. Save the output — analysis, Q\&A answers, and batch summary — to the outputs directory.
**Minimum poll interval: 60 seconds.** The API enforces a 60-second minimum between status calls for the same `job_id`. Do not reduce this value.
***
## Pipeline Implementation
### Imports & Setup
```python imports and config theme={null}
import os, json, time, hashlib, requests
from pathlib import Path
from datetime import datetime
from typing import List, Dict, Optional
from pydub import AudioSegment
try:
import anthropic
except ImportError:
anthropic = None
try:
from openai import OpenAI
except ImportError:
OpenAI = None
OUTPUT_DIR = "outputs"
BATCH_SUBMIT = "https://api.vachana.ai/stt/v3/batch/submit"
BATCH_STATUS = "https://api.vachana.ai/stt/v3/batch/status/{job_id}"
POLL_INTERVAL = 60 # seconds — minimum enforced by the API
LLM_PROVIDER = os.getenv("LLM_PROVIDER", "claude")
Path(OUTPUT_DIR).mkdir(exist_ok=True)
def split_audio(audio_path: str, chunk_ms: int = 3_600_000) -> List[AudioSegment]:
"""Split audio into chunks of at most 1 hour for Batch API compliance."""
audio = AudioSegment.from_file(audio_path)
if len(audio) <= chunk_ms:
return [audio]
return [audio[i:i + chunk_ms] for i in range(0, len(audio), chunk_ms)]
```
### Submit & Poll
```python process_audio_files + _poll_until_complete theme={null}
def process_audio_files(
self,
audio_paths: List[str],
language_code: str = "hi-IN",
itn: bool = True,
) -> Dict[str, dict]:
"""Submit audio files to Vachana Batch STT and poll until complete."""
if not audio_paths:
return {}
files = [
("audio_files", (Path(p).name, open(p, "rb"), "audio/wav"))
for p in audio_paths
]
data = {
"language_code": language_code,
"is_multi_channel": "false",
"format": "transcribe" if itn else "verbatim",
}
resp = requests.post(BATCH_SUBMIT, headers=self.headers, files=files, data=data)
resp.raise_for_status()
job_id = resp.json()["job_id"]
print(f"Job submitted: {job_id}")
for _, (_, fh, _) in files:
fh.close()
results = self._poll_until_complete(job_id)
if not results:
return {}
output_dir = Path(OUTPUT_DIR) / f"job_{job_id}"
output_dir.mkdir(parents=True, exist_ok=True)
transcriptions = self._parse_results(results, output_dir)
self.transcriptions.update(transcriptions)
print(f"Transcribed {len(transcriptions)} file(s).")
for fname, d in transcriptions.items():
self.analyze_transcription(d["conversation_path"], output_dir, fname)
return transcriptions
def _poll_until_complete(self, job_id: str) -> Optional[list]:
"""Poll the status endpoint every 60 s until a terminal state is reached."""
url = BATCH_STATUS.format(job_id=job_id)
print("Polling for results (every 60 s)...")
while True:
time.sleep(POLL_INTERVAL)
r = requests.get(url, headers=self.headers)
r.raise_for_status()
payload = r.json()
status = payload["status"]
print(f" status={status} progress={payload.get('overall_progress', '–')}%")
if status == "completed":
return payload.get("results", [])
if status == "failed":
print(f"Job failed: {payload.get('error')}")
return None
```
### Parsing — Speaker Transcripts & Sentiment Timeline
`_parse_results` processes the completed API response and writes three output files per call: a speaker-labelled transcript, a per-speaker talk-time log, and a segment-level sentiment timeline.
```python _parse_results theme={null}
def _parse_results(self, results: list, output_dir: Path) -> Dict[str, dict]:
"""Parse per-file segment data into conversation and analytics files."""
transcriptions = {}
for file_result in results:
fname = Path(file_result["filename"]).stem
segments = file_result.get("segments", [])
if not segments:
print(f"No segments returned for {fname}, skipping.")
continue
lines, speaker_times, sentiment_log = [], {}, []
for seg in segments:
spk = seg.get("speaker_id", "UNKNOWN")
text = seg.get("text", "").strip()
s = seg.get("start_time", 0.0)
e = seg.get("end_time", 0.0)
lines.append(f"SPEAKER_{spk}: {text}")
speaker_times[spk] = speaker_times.get(spk, 0.0) + (e - s)
sentiment_log.append({
"speaker": spk,
"start_time": s,
"text": text,
"sentiment": seg.get("sentiment", "Neutral"),
"emotion": seg.get("emotion", "Neutral"),
})
conv_path = output_dir / f"{fname}_conversation.txt"
timing_path = output_dir / f"{fname}_timing.json"
sentiment_path = output_dir / f"{fname}_sentiment.json"
conv_path.write_text("\n".join(lines), encoding="utf-8")
timing_path.write_text(json.dumps(speaker_times, indent=2), encoding="utf-8")
sentiment_path.write_text(json.dumps(sentiment_log, indent=2), encoding="utf-8")
transcriptions[fname] = {
"conversation_path": str(conv_path),
"timing_path": str(timing_path),
"sentiment_path": str(sentiment_path),
}
return transcriptions
```
**Files produced per call:** `{name}_conversation.txt` — speaker-labelled transcript · `{name}_timing.json` — talk time per speaker in seconds · `{name}_sentiment.json` — segment-level sentiment and emotion timeline
### LLM Analysis
The analysis step sends the parsed conversation to your chosen LLM with a structured prompt. Switch providers by changing the `LLM_PROVIDER` environment variable.
```python analysis prompt theme={null}
ANALYSIS_PROMPT = """
Analyze this call transcription from start to finish.
TRANSCRIPTION:
{transcription}
Provide a structured response covering each of the following:
1. Speaker identification — which speaker is the customer, which is the agent?
2. Customer type — new/potential customer or existing customer?
3. Opening problem — what issue or query did the customer raise initially?
4. Products or services — what was the customer inquiring about or facing issues with?
5. Agent response — how did the agent handle and resolve the issue throughout the call?
6. Resolution outcome — was the issue resolved? Was the customer satisfied at the end?
7. Sentiment arc — how did the customer's sentiment shift across the call?
8. Upsell or cross-sell signals — any opportunities the agent identified or missed?
9. Competitor mentions — were any competitors referenced?
10. Summary — two-sentence outcome summary.
Note: Segment-level sentiment and emotion are tagged in the transcript where available.
"""
```
```python analyze_transcription + _call_llm theme={null}
def analyze_transcription(self, conversation_path: str, output_dir: Path, file_name: str) -> dict:
"""Run LLM analysis on a parsed conversation file."""
transcript = Path(conversation_path).read_text(encoding="utf-8")
analysis = self._call_llm(
system="You are a call analytics expert. Provide structured, actionable insights.",
user=ANALYSIS_PROMPT.format(transcription=transcript),
)
out = output_dir / f"{file_name}_analysis.txt"
out.write_text(analysis.strip(), encoding="utf-8")
print(f"Analysis saved: {out}")
return {"file_name": file_name, "analysis_path": str(out)}
def _call_llm(self, system: str, user: str) -> str:
"""Route to Claude or OpenAI based on LLM_PROVIDER env variable."""
if LLM_PROVIDER == "claude":
if anthropic is None:
raise ImportError("Install the anthropic package: pip install anthropic")
client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
msg = client.messages.create(
model="claude-opus-4-8",
max_tokens=2000,
system=system,
messages=[{"role": "user", "content": user}],
)
return msg.content[0].text
elif LLM_PROVIDER == "openai":
if OpenAI is None:
raise ImportError("Install the openai package: pip install openai")
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
resp = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": system},
{"role": "user", "content": user},
],
)
return resp.choices[0].message.content
raise ValueError(f"Unknown LLM_PROVIDER '{LLM_PROVIDER}'. Set to 'claude' or 'openai'.")
```
### Ad-hoc Q\&A
Ask any question against a transcribed call — useful for targeted investigation after bulk processing.
```python answer_question theme={null}
def answer_question(self, question: str) -> None:
"""Answer a question for every transcribed call in the current session."""
for fname, data in self.transcriptions.items():
transcript = Path(data["conversation_path"]).read_text(encoding="utf-8")
answer = self._call_llm(
system="",
user=f"TRANSCRIPT:\n{transcript}\n\nQUESTION: {question}",
)
q_hash = hashlib.sha1(question.encode()).hexdigest()[:6]
out = Path(data["conversation_path"]).parent / f"{fname}_q_{q_hash}.txt"
out.write_text(f"Q: {question}\n\nA:\n{answer}", encoding="utf-8")
print(f"Answer saved: {out}")
```
### Summary Report
Generate a single summary report across all analyzed calls in the session.
```python get_summary theme={null}
SUMMARY_PROMPT = """
Based on this call analysis, provide a concise 2–3 word answer for each point:
{analysis_text}
1. Customer and Agent
2. Customer Type
3. Main Issue
4. Service Discussed
5. Agent Response Quality
6. Customer Satisfaction
7. Overall Sentiment
8. Competitor or Upsell Signal
9. Resolution Status
"""
def get_summary(self) -> None:
"""Generate a single summary report across all calls in the session."""
ts = datetime.now().strftime("%Y%m%d_%H%M%S")
out = Path(OUTPUT_DIR) / f"summary_{ts}.txt"
with open(out, "w", encoding="utf-8") as f:
f.write(f"CALL ANALYTICS SUMMARY\n{'='*60}\n")
f.write(f"Generated : {datetime.now()}\n")
f.write(f"Total calls: {len(self.transcriptions)}\n{'='*60}\n\n")
for fname, data in self.transcriptions.items():
af = Path(data["conversation_path"]).parent / f"{fname}_analysis.txt"
if not af.exists():
print(f"No analysis file found for {fname}, skipping.")
continue
summary = self._call_llm(
system="You are a call analytics expert. Be concise.",
user=SUMMARY_PROMPT.format(analysis_text=af.read_text(encoding="utf-8")),
)
f.write(f"Call: {fname}\n{'-'*30}\n{summary.strip()}\n\n")
print(f"Summary saved: {out}")
```
***
## Full Pipeline
```python call_analytics_pipeline.py theme={null}
import os, json, time, hashlib, requests
from pathlib import Path
from datetime import datetime
from typing import List, Dict, Optional
from pydub import AudioSegment
try:
import anthropic
except ImportError:
anthropic = None
try:
from openai import OpenAI
except ImportError:
OpenAI = None
OUTPUT_DIR = "outputs"
BATCH_SUBMIT = "https://api.vachana.ai/stt/v3/batch/submit"
BATCH_STATUS = "https://api.vachana.ai/stt/v3/batch/status/{job_id}"
POLL_INTERVAL = 60
LLM_PROVIDER = os.getenv("LLM_PROVIDER", "claude")
Path(OUTPUT_DIR).mkdir(exist_ok=True)
ANALYSIS_PROMPT = """
Analyze this call transcription from start to finish.
TRANSCRIPTION:
{transcription}
Provide a structured response covering each of the following:
1. Speaker identification — which speaker is the customer, which is the agent?
2. Customer type — new/potential customer or existing customer?
3. Opening problem — what issue or query did the customer raise initially?
4. Products or services — what was the customer inquiring about or facing issues with?
5. Agent response — how did the agent handle and resolve the issue throughout the call?
6. Resolution outcome — was the issue resolved? Was the customer satisfied at the end?
7. Sentiment arc — how did the customer's sentiment shift across the call?
8. Upsell or cross-sell signals — any opportunities the agent identified or missed?
9. Competitor mentions — were any competitors referenced?
10. Summary — two-sentence outcome summary.
"""
SUMMARY_PROMPT = """
Based on this call analysis, provide a concise 2–3 word answer for each point:
{analysis_text}
1. Customer and Agent
2. Customer Type
3. Main Issue
4. Service Discussed
5. Agent Response Quality
6. Customer Satisfaction
7. Overall Sentiment
8. Competitor or Upsell Signal
9. Resolution Status
"""
def split_audio(audio_path: str, chunk_ms: int = 3_600_000) -> List[AudioSegment]:
audio = AudioSegment.from_file(audio_path)
if len(audio) <= chunk_ms:
return [audio]
return [audio[i:i + chunk_ms] for i in range(0, len(audio), chunk_ms)]
class CallAnalytics:
def __init__(self, api_key: str):
self.headers = {"X-API-Key-ID": api_key}
self.transcriptions: Dict[str, dict] = {}
def process_audio_files(self, audio_paths: List[str], language_code: str = "hi-IN", itn: bool = True) -> Dict[str, dict]:
if not audio_paths:
return {}
files = [("audio_files", (Path(p).name, open(p, "rb"), "audio/wav")) for p in audio_paths]
data = {"language_code": language_code, "is_multi_channel": "false",
"format": "transcribe" if itn else "verbatim"}
resp = requests.post(BATCH_SUBMIT, headers=self.headers, files=files, data=data)
resp.raise_for_status()
job_id = resp.json()["job_id"]
print(f"Job submitted: {job_id}")
for _, (_, fh, _) in files: fh.close()
results = self._poll_until_complete(job_id)
if not results: return {}
output_dir = Path(OUTPUT_DIR) / f"job_{job_id}"
output_dir.mkdir(parents=True, exist_ok=True)
transcriptions = self._parse_results(results, output_dir)
self.transcriptions.update(transcriptions)
print(f"Transcribed {len(transcriptions)} file(s).")
for fname, d in transcriptions.items():
self.analyze_transcription(d["conversation_path"], output_dir, fname)
return transcriptions
def _poll_until_complete(self, job_id: str) -> Optional[list]:
url = BATCH_STATUS.format(job_id=job_id)
print("Polling every 60 s...")
while True:
time.sleep(POLL_INTERVAL)
r = requests.get(url, headers=self.headers)
r.raise_for_status()
p = r.json()
print(f" [{p['status']}] {p.get('overall_progress', '–')}%")
if p["status"] == "completed": return p.get("results", [])
if p["status"] == "failed":
print(f"Job failed: {p.get('error')}")
return None
def _parse_results(self, results: list, output_dir: Path) -> Dict[str, dict]:
transcriptions = {}
for fr in results:
fname = Path(fr["filename"]).stem
segments = fr.get("segments", [])
if not segments: continue
lines, speaker_times, sentiment_log = [], {}, []
for seg in segments:
spk = seg.get("speaker_id", "UNKNOWN")
txt = seg.get("text", "").strip()
s, e = seg.get("start_time", 0.0), seg.get("end_time", 0.0)
lines.append(f"SPEAKER_{spk}: {txt}")
speaker_times[spk] = speaker_times.get(spk, 0.0) + (e - s)
sentiment_log.append({"speaker": spk, "start_time": s, "text": txt,
"sentiment": seg.get("sentiment", "Neutral"),
"emotion": seg.get("emotion", "Neutral")})
cp = output_dir / f"{fname}_conversation.txt"
tp = output_dir / f"{fname}_timing.json"
sp = output_dir / f"{fname}_sentiment.json"
cp.write_text("\n".join(lines), encoding="utf-8")
tp.write_text(json.dumps(speaker_times, indent=2), encoding="utf-8")
sp.write_text(json.dumps(sentiment_log, indent=2), encoding="utf-8")
transcriptions[fname] = {"conversation_path": str(cp), "timing_path": str(tp), "sentiment_path": str(sp)}
return transcriptions
def _call_llm(self, system: str, user: str) -> str:
if LLM_PROVIDER == "claude":
client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
msg = client.messages.create(model="claude-opus-4-8", max_tokens=2000,
system=system, messages=[{"role": "user", "content": user}])
return msg.content[0].text
elif LLM_PROVIDER == "openai":
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
resp = client.chat.completions.create(model="gpt-4o",
messages=[{"role": "system", "content": system},
{"role": "user", "content": user}])
return resp.choices[0].message.content
raise ValueError(f"Unknown LLM_PROVIDER: {LLM_PROVIDER}")
def analyze_transcription(self, conversation_path: str, output_dir: Path, fname: str):
transcript = Path(conversation_path).read_text(encoding="utf-8")
analysis = self._call_llm(system="You are a call analytics expert. Provide structured, actionable insights.",
user=ANALYSIS_PROMPT.format(transcription=transcript))
out = output_dir / f"{fname}_analysis.txt"
out.write_text(analysis.strip(), encoding="utf-8")
print(f"Analysis: {out}")
def answer_question(self, question: str):
for fname, data in self.transcriptions.items():
transcript = Path(data["conversation_path"]).read_text(encoding="utf-8")
answer = self._call_llm(system="",
user=f"TRANSCRIPT:\n{transcript}\n\nQUESTION: {question}")
q_hash = hashlib.sha1(question.encode()).hexdigest()[:6]
out = Path(data["conversation_path"]).parent / f"{fname}_q_{q_hash}.txt"
out.write_text(f"Q: {question}\n\nA:\n{answer}", encoding="utf-8")
print(f"Q&A: {out}")
def get_summary(self):
ts = datetime.now().strftime("%Y%m%d_%H%M%S")
out = Path(OUTPUT_DIR) / f"summary_{ts}.txt"
with open(out, "w", encoding="utf-8") as f:
f.write(f"CALL ANALYTICS SUMMARY\n{'='*60}\nGenerated: {datetime.now()}\n{'='*60}\n\n")
for fname, data in self.transcriptions.items():
af = Path(data["conversation_path"]).parent / f"{fname}_analysis.txt"
if not af.exists(): continue
summary = self._call_llm(system="Be concise.",
user=SUMMARY_PROMPT.format(analysis_text=af.read_text(encoding="utf-8")))
f.write(f"Call: {fname}\n{'-'*30}\n{summary.strip()}\n\n")
print(f"Summary: {out}")
if __name__ == "__main__":
analytics = CallAnalytics(api_key=os.getenv("GNANI_API_KEY"))
analytics.process_audio_files(
audio_paths=["./call_001.wav"],
language_code="hi-IN",
itn=True,
)
analytics.answer_question("Did the agent offer any EMI or payment extension options?")
analytics.get_summary()
```
***
## Sample Output
```text theme={null}
outputs/
└── job_batch_7f3a92c1d4e8/
├── call_001_conversation.txt ← speaker-labelled transcript
├── call_001_timing.json ← talk time per speaker (seconds)
├── call_001_sentiment.json ← segment-level sentiment timeline
├── call_001_analysis.txt ← LLM structured analysis
└── call_001_q_a3f9b2.txt ← ad-hoc Q&A answer
summary_20251226_143052.txt ← batch summary across all calls
```
**call\_001\_conversation.txt**
```text theme={null}
SPEAKER_1: नमस्ते, मेरा नाम रोहन है। मेरी EMI अगले हफ्ते due है।
SPEAKER_2: नमस्ते रोहन जी, आपका loan account number बताइए।
SPEAKER_1: हाँ, ₹45,000 की EMI है। क्या मुझे extension मिल सकता है?
SPEAKER_2: आपकी request process करते हैं। 3 दिन का extension approve हो सकता है।
SPEAKER_1: ठीक है, शुक्रिया।
```
**call\_001\_analysis.txt (excerpt)**
```text theme={null}
1. Speaker Identification
SPEAKER_1 — Customer (Rohan)
SPEAKER_2 — Agent
2. Customer Type
Existing customer with an active loan account and an upcoming EMI.
3. Opening Problem
Customer called to request an EMI payment extension due to cash flow constraints.
6. Resolution Outcome
Resolved within the call. Customer expressed satisfaction before closing.
7. Sentiment Arc
Started neutral-to-anxious. Shifted to relieved after the extension was confirmed.
8. Upsell / Cross-sell Signals
No signals identified or pursued. Loan restructuring or a credit health check
could have been offered — it was not.
10. Summary
The customer's EMI extension request was resolved within a single call.
Agent resolution quality was high; a potential upsell moment was missed.
```
***
## Limits & Notes
| Constraint | Value | What to do |
| ------------------------- | ------------------ | ---------------------------------------------------------------------------------------- |
| **Max file duration** | 1 hour per file | Use `split_audio()` for longer recordings. Concatenate conversation text after parsing. |
| **Max files per request** | 10 files | For bulk pipelines, group calls into batches of 10 and submit sequentially. |
| **Max payload size** | 80 MB total | Use FLAC or Opus to reduce file size before upload. |
| **Poll interval** | 60 seconds minimum | Do not reduce below 60 s. The API rate-limits per `job_id`. |
| **Speaker diarization** | Max 2 speakers | Designed for two-party calls (agent + customer). Multi-party calls are not supported. |
| **ITN support** | hi-IN, en-IN only | Other languages return verbatim output regardless of the `format` parameter. |
| **LLM token limits** | Varies by provider | For calls over 30 minutes, chunk the transcript before sending to the LLM analysis step. |
**Related docs:** [Batch STT API reference](/vachana/STT/stt-batch) · [REST STT for short clips](/vachana/STT/speech-to-text) · SDK: `pip install gnani-vachana`
# Podcast Transcription with Speaker Labels
Source: https://docs.gnani.ai/api/use-cases/podcast-transcription
Transcribe multi-speaker audio at scale using the Gnani Prisma v2.5 Batch STT API. From submitting an audio file to receiving a clean, speaker-separated transcript — in 10 Indian languages.
## Overview
Transcribe multi-speaker audio at scale using the Gnani Prisma v2.5 Batch STT API. This guide walks through every step — from submitting an audio file to receiving a clean, speaker-separated transcript — using podcast transcription as the working example.
Audio-first content — podcasts, interview recordings, panel discussions — carries information that stays locked unless it is transcribed. Speaker-level transcription is what separates a readable document from a wall of undifferentiated text. You know who said what, when they said it, and for how long.
| Capability | What it enables downstream |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Speaker-separated output** | Per-speaker text blocks mean editors can review one voice at a time, and content teams can attribute quotes accurately. |
| **Time-aligned segments** | Every segment carries a `start_time` and `end_time`, enabling subtitle generation, chapter markers, and clip extraction at a specific timestamp. |
| **Segment-level confidence** | Flag low-confidence segments for human review rather than reviewing the entire transcript. |
| **10 Indian languages** | Transcribe Hindi, Tamil, Telugu, Kannada, Malayalam, Bengali, Gujarati, Marathi, Punjabi, and English without switching providers or pipelines. |
| **Batch processing** | Submit up to 10 files in a single API call. Run overnight jobs, backfill archives, or process weekly episode batches without managing queues yourself. |
***
## Other Use Cases
The same submit-poll-parse pipeline works for any long-form, speaker-rich audio. Any scenario involving long audio files, two speakers, and a need for speaker-separated text maps directly to this pipeline.
| Use Case | Description |
| ----------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Journalist Interviews** | Transcribe field recordings with interviewer and subject separated. Feed directly into editorial workflows without manual formatting. |
| **Parliamentary & Panel Debates** | Attribute statements to the correct speaker for political reporting, fact-checking, or archival. Supports Devanagari and regional scripts natively. |
| **EdTech Lecture Recordings** | Transcribe faculty and student exchanges. Generate accessible transcripts for students, search indexes for course platforms, and study material exports. |
| **Legal Depositions & Hearings** | Produce verbatim speaker-attributed records of proceedings. Use confidence scores to flag segments requiring court reporter verification. |
| **Radio Archive Digitisation** | Backfill years of archived broadcasts into searchable, attributed text. Batch processing handles large volumes without manual queuing. |
| **Corporate Town Halls & Earnings Calls** | Generate attributed transcripts of leadership Q\&A sessions. Surface speaker-specific statements for internal comms or investor relations. |
| **Documentary & Film Production** | Auto-generate interview transcripts for rough-cut editing. Export time-coded speaker lines directly to editing software. |
| **Doctor-Patient Consultations** | Transcribe recorded consultations with doctor and patient separated. Enable structured documentation workflows for EMR systems. |
**Two-speaker limit:** The Gnani Prisma v2.5 Batch STT API supports a maximum of two distinct speakers per file. It is optimised for two-party audio — interviews, conversations, and one-on-one recordings. Panel discussions with three or more speakers are outside the current scope.
***
## Prerequisites
| Requirement | Details |
| ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Gnani Prisma v2.5 API key** | Available from the Gnani API dashboard. You will use this as the `X-API-Key-ID` header on every request. |
| **Python 3.9+** | The pipeline uses f-strings, `pathlib`, and `typing` patterns that require Python 3.9 or later. |
| **Audio files** | Supported formats: AAC, WAV, FLAC, ALAC, OGG (Vorbis), Opus. Each file must be under 1 hour and the total payload under 80 MB. |
| **ffmpeg** | Required only if you plan to split files longer than 1 hour using pydub. Install with `brew install ffmpeg` (macOS) or `apt install ffmpeg` (Linux). |
```bash theme={null}
# HTTP client (used for submit and poll calls)
pip install requests
# Audio chunking — only needed for files over 1 hour
pip install pydub
```
No SDK is required for this pipeline. All calls use the standard HTTP REST endpoints.
***
## Authentication
| Header | Required | Description |
| ------------------ | -------- | -------------------------------------------------------------------------------------------------- |
| `X-API-Key-ID` | Yes | Your Gnani Prisma v2.5 API key. Required on both the submit and status calls. |
| `X-API-Request-ID` | No | A UUID you assign for tracing. Useful for correlating your application logs with platform support. |
Store your API key as an environment variable. Never hardcode it in source files or commit it to version control.
```bash .env theme={null}
GNANI_API_KEY=your-api-key-here
```
```python loading credentials theme={null}
import os
API_KEY = os.getenv("GNANI_API_KEY")
HEADERS = {"X-API-Key-ID": API_KEY}
```
**Never hardcode API keys.** Do not commit credentials to version control. Use environment variables, a secrets manager, or a vault. Rotate your key immediately if it is exposed.
***
## Limits & Supported Formats
| Item | Limit |
| ---------------------- | ----------------------------------------------------- |
| Max audio duration | Less than 1 hour per file |
| Max files per request | 10 files per API call |
| Max total payload size | 80 MB across all files and form fields combined |
| Minimum poll interval | 60 seconds between status calls for the same `job_id` |
| Speaker diarization | Maximum 2 speakers per file |
### Supported Audio Formats
| Format | Extension | Notes |
| ---------------- | --------- | ------------------------------------------------------------------------------------------ |
| **WAV** | `.wav` | Uncompressed. Highest quality but largest file size. |
| **FLAC** | `.flac` | Lossless compression. Good balance of quality and size for archival audio. |
| **AAC** | `.m4a` | Common podcast export format. Well-supported across all recording tools. |
| **ALAC** | `.m4a` | Lossless Apple format. Use when source is from Apple recording tools. |
| **OGG (Vorbis)** | `.ogg` | Open format. Common in Linux recording pipelines. |
| **Opus** | `.opus` | Efficient lossy compression. Smallest file sizes — recommended for high-volume batch jobs. |
**Files over 1 hour:** Split into chunks before submitting. The `split_audio()` helper in the full script handles this automatically using pydub. Stitch the resulting transcripts in order after parsing.
***
## Supported Languages
Pass the BCP-47 code in the `language_code` field of your submit request.
| Language | Code | Native Script | ITN |
| ------------- | ------- | ------------- | --- |
| **Bengali** | `bn-IN` | বাংলা | — |
| **English** | `en-IN` | Latin | Yes |
| **Gujarati** | `gu-IN` | ગુજરાતી | — |
| **Hindi** | `hi-IN` | हिन्दी | Yes |
| **Kannada** | `kn-IN` | ಕನ್ನಡ | — |
| **Malayalam** | `ml-IN` | മലയാളം | — |
| **Marathi** | `mr-IN` | मराठी | — |
| **Punjabi** | `pa-IN` | ਪੰਜਾਬੀ | — |
| **Tamil** | `ta-IN` | தமிழ் | — |
| **Telugu** | `te-IN` | తెలుగు | — |
ITN (Inverse Text Normalization) converts spoken numbers, currency, dates, times, and phone numbers into compact written form. Currently available for `hi-IN` and `en-IN` only. Set `format=transcribe` in the request to enable it.
***
## Pipeline
POST to `/stt/v3/batch/submit` with your audio file, language code, and format preference. Receive a `job_id` immediately. Save it — you will need it in every poll call.
GET `/stt/v3/batch/status/{job_id}` every 60 seconds. Loop until `status` reaches `completed` or `failed`. The `results` field is `null` until the job is complete.
Iterate over `results[].segments`. Group by `speaker_id`. Build per-speaker text blocks with timestamps and confidence scores.
Write the speaker-labelled transcript to a text file, along with a JSON file containing per-speaker talk time and segment metadata.
### Step 1 — Submit
```python submit_job() theme={null}
import os
import requests
from pathlib import Path
BATCH_SUBMIT = "https://api.vachana.ai/stt/v3/batch/submit"
def submit_job(
audio_path: str,
language_code: str = "hi-IN",
itn: bool = True,
) -> str:
api_key = os.getenv("GNANI_API_KEY")
headers = {"X-API-Key-ID": api_key}
with open(audio_path, "rb") as f:
files = [("audio_files", (Path(audio_path).name, f, "audio/wav"))]
data = {
"language_code": language_code,
"is_multi_channel": "false",
"format": "transcribe" if itn else "verbatim",
}
resp = requests.post(BATCH_SUBMIT, headers=headers, files=files, data=data)
resp.raise_for_status()
job_id = resp.json()["job_id"]
print(f"Submitted. job_id: {job_id}")
return job_id
```
**Submitting multiple files:** Pass additional `("audio_files", ...)` tuples to the `files` list. Up to 10 files are accepted per request, as long as the total payload stays under 80 MB.
### Step 2 — Poll
```python poll_until_complete() theme={null}
import time
from typing import Optional
BATCH_STATUS = "https://api.vachana.ai/stt/v3/batch/status/{job_id}"
POLL_INTERVAL = 60 # seconds — enforced minimum; do not reduce
def poll_until_complete(job_id: str) -> Optional[list]:
api_key = os.getenv("GNANI_API_KEY")
headers = {"X-API-Key-ID": api_key}
url = BATCH_STATUS.format(job_id=job_id)
print(f"Polling job {job_id} every {POLL_INTERVAL}s...")
while True:
time.sleep(POLL_INTERVAL)
resp = requests.get(url, headers=headers)
resp.raise_for_status()
payload = resp.json()
status = payload["status"]
progress = payload.get("overall_progress", "–")
print(f" [{status}] progress: {progress}%")
if status == "completed":
print(f"Job complete. {payload['completed_files']} file(s) transcribed.")
return payload.get("results", [])
if status == "failed":
print(f"Job failed: {payload.get('error')}")
return None
```
**Minimum poll interval: 60 seconds.** The API enforces a rate limit of one status call per 60 seconds per `job_id`. Do not reduce it.
### Step 3 — Parse
```python parse_results() theme={null}
import json
from pathlib import Path
from typing import Dict
def parse_results(results: list, output_dir: Path) -> Dict[str, dict]:
outputs = {}
for file_result in results:
fname = Path(file_result["filename"]).stem
segments = file_result.get("segments", [])
if not segments or file_result.get("status") == "failed":
print(f"Skipping {fname}: {file_result.get('error', 'no segments')}")
continue
lines, speaker_times, segment_meta = [], {}, []
for seg in segments:
spk = seg.get("speaker_id", "UNKNOWN")
text = seg.get("text", "").strip()
start = seg.get("start_time", 0.0)
end = seg.get("end_time", 0.0)
ts = f"{int(start // 60):02d}:{int(start % 60):02d}"
lines.append(f"[{ts}] SPEAKER_{spk}: {text}")
speaker_times[spk] = speaker_times.get(spk, 0.0) + (end - start)
segment_meta.append({
"segment_id": seg.get("segment_id"),
"speaker_id": spk,
"start_time": start,
"end_time": end,
"text": text,
"confidence": seg.get("confidence"),
"language_detected": seg.get("language_detected"),
})
transcript_path = output_dir / f"{fname}_transcript.txt"
transcript_path.write_text("\n".join(lines), encoding="utf-8")
metadata_path = output_dir / f"{fname}_metadata.json"
metadata_path.write_text(json.dumps({
"filename": file_result["filename"],
"total_duration_s": file_result.get("total_duration"),
"speaker_talk_time": {f"SPEAKER_{k}": round(v, 2) for k, v in speaker_times.items()},
"segments": segment_meta,
}, indent=2, ensure_ascii=False), encoding="utf-8")
outputs[fname] = {
"transcript_path": str(transcript_path),
"metadata_path": str(metadata_path),
}
print(f"Parsed: {fname} → {len(lines)} segments, {len(speaker_times)} speaker(s)")
return outputs
```
***
## Full Script
```python podcast_transcription.py theme={null}
import os
import json
import time
import requests
from pathlib import Path
from typing import Dict, List, Optional
BATCH_SUBMIT = "https://api.vachana.ai/stt/v3/batch/submit"
BATCH_STATUS = "https://api.vachana.ai/stt/v3/batch/status/{job_id}"
POLL_INTERVAL = 60
OUTPUT_DIR = "outputs"
Path(OUTPUT_DIR).mkdir(exist_ok=True)
def submit_job(audio_paths: List[str], language_code: str = "hi-IN", itn: bool = True) -> str:
api_key = os.getenv("GNANI_API_KEY")
headers = {"X-API-Key-ID": api_key}
files = [("audio_files", (Path(p).name, open(p, "rb"), "audio/wav")) for p in audio_paths]
data = {"language_code": language_code, "is_multi_channel": "false",
"format": "transcribe" if itn else "verbatim"}
resp = requests.post(BATCH_SUBMIT, headers=headers, files=files, data=data)
resp.raise_for_status()
for _, (_, fh, _) in files:
fh.close()
job_id = resp.json()["job_id"]
print(f"Submitted {len(audio_paths)} file(s). job_id: {job_id}")
return job_id
def poll_until_complete(job_id: str) -> Optional[list]:
api_key = os.getenv("GNANI_API_KEY")
headers = {"X-API-Key-ID": api_key}
url = BATCH_STATUS.format(job_id=job_id)
print(f"Polling every {POLL_INTERVAL}s...")
while True:
time.sleep(POLL_INTERVAL)
resp = requests.get(url, headers=headers)
resp.raise_for_status()
payload = resp.json()
status = payload["status"]
print(f" [{status}] {payload.get('overall_progress', '–')}%")
if status == "completed":
return payload.get("results", [])
if status == "failed":
print(f"Job failed: {payload.get('error')}")
return None
def parse_results(results: list, output_dir: Path) -> Dict[str, dict]:
outputs = {}
for file_result in results:
fname = Path(file_result["filename"]).stem
segments = file_result.get("segments", [])
if not segments or file_result.get("status") == "failed":
print(f"Skipping {fname}: {file_result.get('error', 'no segments')}")
continue
lines, speaker_times, segment_meta = [], {}, []
for seg in segments:
spk = seg.get("speaker_id", "UNKNOWN")
text = seg.get("text", "").strip()
start = seg.get("start_time", 0.0)
end = seg.get("end_time", 0.0)
ts = f"{int(start // 60):02d}:{int(start % 60):02d}"
lines.append(f"[{ts}] SPEAKER_{spk}: {text}")
speaker_times[spk] = speaker_times.get(spk, 0.0) + (end - start)
segment_meta.append({
"segment_id": seg.get("segment_id"),
"speaker_id": spk,
"start_time": start,
"end_time": end,
"text": text,
"confidence": seg.get("confidence"),
"language_detected": seg.get("language_detected"),
})
transcript_path = output_dir / f"{fname}_transcript.txt"
transcript_path.write_text("\n".join(lines), encoding="utf-8")
metadata_path = output_dir / f"{fname}_metadata.json"
metadata_path.write_text(json.dumps({
"filename": file_result["filename"],
"total_duration_s": file_result.get("total_duration"),
"speaker_talk_time": {f"SPEAKER_{k}": round(v, 2) for k, v in speaker_times.items()},
"segments": segment_meta,
}, indent=2, ensure_ascii=False), encoding="utf-8")
outputs[fname] = {"transcript_path": str(transcript_path), "metadata_path": str(metadata_path)}
print(f"Saved: {transcript_path.name}")
return outputs
if __name__ == "__main__":
job_id = submit_job(audio_paths=["/path/to/episode_01.wav"], language_code="hi-IN", itn=True)
results = poll_until_complete(job_id)
if results:
output_dir = Path(OUTPUT_DIR) / f"job_{job_id}"
output_dir.mkdir(parents=True, exist_ok=True)
outputs = parse_results(results, output_dir)
print(f"\nDone. {len(outputs)} transcript(s) saved to {output_dir}/")
```
***
## Sample Output
```text theme={null}
outputs/
└── job_batch_7f3a92c1d4e8/
├── episode_01_transcript.txt ← speaker-labelled, time-stamped transcript
└── episode_01_metadata.json ← duration, talk time, segment detail
```
**episode\_01\_transcript.txt**
```text theme={null}
[00:00] SPEAKER_1: नमस्ते, मैं हूँ रवि शर्मा और आज हम बात करेंगे भारत के स्टार्टअप इकोसिस्टम के बारे में।
[00:07] SPEAKER_2: हाँ रवि जी, बहुत अच्छा विषय है। पिछले पाँच साल में बहुत कुछ बदला है।
[00:14] SPEAKER_1: बिल्कुल। ₹2,00,000 करोड़ से ज़्यादा की फंडिंग आई है 2024 में।
[00:22] SPEAKER_2: और यूनिकॉर्न्स की संख्या भी 100 के पार पहुँच गई है।
```
**episode\_01\_metadata.json**
```json theme={null}
{
"filename": "episode_01.wav",
"total_duration_s": 2847.5,
"speaker_talk_time": {
"SPEAKER_1": 1423.8,
"SPEAKER_2": 1389.2
},
"segments": [
{
"segment_id": 0,
"speaker_id": 1,
"start_time": 0.0,
"end_time": 6.8,
"text": "नमस्ते, मैं हूँ रवि शर्मा...",
"confidence": 0.96,
"language_detected": "hi-IN"
}
]
}
```
***
## ITN Reference
When `format=transcribe` is set on the submit request, ITN post-processes every transcript — converting spoken-form numbers, currency, dates, times, and phone numbers into compact written form. Available for `hi-IN` and `en-IN` only.
| Category | Spoken input (ASR) | Written output (ITN) |
| ----------------- | ------------------------------------- | --------------------------- |
| **Numbers** | पाँच लाख बीस हज़ार | 5,20,000 |
| **Currency** | तीन रुपये पचास पैसे | ₹3.50 |
| **Currency (en)** | five thousand rupees | ₹5,000 |
| **Dates** | बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 |
| **Times** | शाम पाँच बजे | शाम 17:00 |
| **Phone numbers** | नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 |
| **Code-mixed** | pay do lakh rupees by fifteenth march | pay ₹2,00,000 by 15th March |
**Native script digits:** Pass `itn_native_numerals=true` alongside `format=transcribe` to render digits in the native script of the target language — for example, Devanagari numerals (₹५,०००) for Hindi. English always outputs Western Arabic digits regardless of this setting.
**What ITN does not change:** Idiomatic and ambiguous phrases are preserved intentionally. `दो तीन` meaning "a few" stays as text, not `2` or `3`. Imperative verbs like `कर दो` or `ले दो` are kept as words.
# Real-Time Quality & Compliance Monitoring
Source: https://docs.gnani.ai/api/use-cases/real-time-compliance
Build a production-grade monitoring system that streams contact center audio to the Gnani Prisma v2.5 WebSocket STT API, detects compliance violations and quality signals in live transcripts, and triggers alerts in under 200ms of speech completion.
## Overview
Contact centers handling financial services, insurance, or healthcare operate under strict regulatory requirements. Agents must follow scripts, disclose specific information, and avoid prohibited language. Traditional QA reviews 2–5% of calls after the fact. By the time a violation is caught, it has already happened hundreds of times.
This guide shows you how to build a system that monitors every call in real time. Audio streams to the Gnani Prisma v2.5 WebSocket STT API. Transcripts arrive within milliseconds of speech completion. A compliance and quality engine processes each segment, matches against rule sets, and fires alerts to your backend — while the call is still live.
| Capability | Implementation |
| ------------------------ | ------------------------------------------------------------------------------------------- |
| **Live transcription** | WebSocket stream to `wss://api.vachana.ai/stt/v3/stream` with per-segment transcript events |
| **Compliance detection** | Keyword and phrase matching on each `transcript` event with configurable rule sets |
| **Quality monitoring** | Silence detection, interruption tracking, escalation phrase matching from segment metadata |
| **Real-time alerts** | Async alert dispatcher — webhook, queue, or supervisor dashboard |
| **Reconnect handling** | Exponential backoff with session continuity across drops |
**Which API to use?** This use case uses the **WebSocket STT API** for real-time streaming. For post-call batch analysis, see the [Call Analytics Pipeline](/vachana/use-cases/call-analytics) which uses the Batch STT API.
***
## Architecture
The system has three logical layers: audio ingestion, transcription, and monitoring. Each runs concurrently in an async event loop.
```text theme={null}
AUDIO SOURCE
│ (Telephony bridge / RTP tap / microphone)
│ PCM 16-bit LE, 16kHz or 8kHz, mono
↓
AUDIO STREAMER
│ Chunks audio into 1024-byte frames (32ms @ 16kHz)
│ Maintains real-time cadence — no burst, no starvation
↓
VACHANA WEBSOCKET STT API wss://api.vachana.ai/stt/v3/stream
│ VAD detects speech boundaries
│ Returns: connected → processing → transcript events
│ Latency: ~300–500ms from end of speech to transcript
↓
TRANSCRIPT HANDLER → COMPLIANCE ENGINE
→ QUALITY ENGINE
↓
ALERT DISPATCHER → Webhook / Queue / Supervisor dashboard
```
Each call owns an isolated **session object** that tracks the full transcript buffer, a timeline of events, compliance status, quality metrics, and reconnect context. This state survives WebSocket reconnects and is flushed to your store at call end.
***
## Prerequisites
| Requirement | Details |
| ----------------- | ------------------------------------------------------------------------------------------------------ |
| **Gnani API key** | Available from the Gnani API dashboard. Used as the `x-api-key-id` header on the WebSocket connection. |
| **Python 3.9+** | Required by the SDK. The full example uses `asyncio`, `dataclasses`, and typed event classes. |
| **Audio source** | PCM 16-bit LE, mono. Either 8kHz (PSTN/legacy VoIP) or 16kHz (wideband VoIP). Defaults to 16kHz. |
| **Alert target** | An HTTP endpoint, message queue, or Redis channel to receive alerts. |
```bash theme={null}
pip install gnani-vachana aiohttp python-dotenv
```
***
## Authentication
Authentication is performed at connection time via HTTP headers on the WebSocket upgrade request. There is no separate auth step — the connection either opens or returns 401.
| Header | Required | Description |
| --------------- | -------- | ------------------------------------------------------------------------------------------------------------- |
| `x-api-key-id` | Yes | Your Gnani Prisma v2.5 API key. |
| `lang_code` | Yes | BCP-47 language code. Defaults to `en-IN`. Pass comma-separated codes for multilingual auto-detection. |
| `x-sample-rate` | No | Audio sample rate in Hz. Accepted: `8000`, `16000`, `44100`, `48000`. Defaults to `16000`. |
| `x-format` | No | Set `transcribe` for ITN (numbers, currency, dates in written form). ITN applies to `hi-IN` and `en-IN` only. |
```bash .env theme={null}
GNANI_API_KEY=your-api-key-here
ALERT_WEBHOOK_URL=https://supervisor.internal/alerts
LANG_CODE=hi-IN
SAMPLE_RATE=16000
```
**Never hardcode API keys.** Load credentials from environment variables or a secrets manager. The `x-api-key-id` header is visible in plaintext in WebSocket upgrade logs — ensure those logs are access-controlled.
***
## End-to-End Workflow
Your telephony bridge fires a call-start event. The monitor opens a WebSocket to `wss://api.vachana.ai/stt/v3/stream` with auth headers and language config. A session object is created and keyed to the call ID.
The server returns a `connected` event confirming sample rate and chunk size. Any mismatch (wrong sample rate, unsupported language) surfaces immediately.
An async producer task reads PCM frames from the telephony tap and sends them at real-time cadence: one 1024-byte frame every 32ms for 16kHz audio. Bursting frames degrades VAD accuracy.
When VAD detects end-of-speech, the server sends a `processing` event. Use this timestamp to measure speech-to-transcript latency and to start a silence timer in the quality engine.
The `transcript` event carries `text`, `segment_index`, `audio_duration_ms`, and `latency`. Both engines process the text synchronously. Alerts are dispatched async so they never block the next transcript.
Compliance violations and quality alerts go to the alert dispatcher. Severity determines the channel: `CRITICAL` hits the supervisor dashboard immediately; `WARNING` queues for post-call review.
On call end, close the WebSocket gracefully. Run final session-level checks (e.g. required disclosure was never spoken). Flush session state to your store and emit a call-complete summary event.
***
## Connecting to the WebSocket API
The SDK's `GnaniSTTStreamClient` wraps the WebSocket connection, frame pacing, and event parsing. Use it as an async context manager.
```python basic connection theme={null}
import asyncio, os
from gnani.stt import GnaniSTTStreamClient
async def open_stream():
async with GnaniSTTStreamClient(
api_key=os.getenv("GNANI_API_KEY"),
language_code="hi-IN", # or comma-separated for auto-detect
sample_rate=16000,
) as stream:
async for event in stream:
await handle_event(event)
```
**Multilingual auto-detection:** For multilingual contact centers, pass comma-separated codes as `lang_code` (e.g. `hi-IN,ta-IN,en-IN`). The API detects the dominant language per segment. Adds minimal latency but removes the need to pre-classify calls by language.
***
## Streaming Audio
### Audio format requirements
| Property | 16kHz (wideband VoIP) | 8kHz (PSTN / legacy) |
| ----------------- | ------------------------------- | ------------------------------- |
| **Encoding** | PCM signed 16-bit little-endian | PCM signed 16-bit little-endian |
| **Channels** | 1 (mono) | 1 (mono) |
| **Frame size** | 1024 bytes (512 samples = 32ms) | 1024 bytes (512 samples = 64ms) |
| **x-sample-rate** | `16000` | `8000` |
Each WebSocket frame must be exactly **1024 bytes**. Bursting frames (sending faster than real time) degrades VAD accuracy — the VAD model is trained on real-time cadence.
```python audio producer task theme={null}
import asyncio
FRAME_SIZE = 1024 # bytes — exactly 512 x 16-bit samples
FRAME_MS_16K = 0.032 # 32ms per frame at 16kHz
FRAME_MS_8K = 0.064 # 64ms per frame at 8kHz
async def stream_audio_producer(stream, audio_source, sample_rate=16000, stop_event=None):
frame_interval = FRAME_MS_16K if sample_rate == 16000 else FRAME_MS_8K
buffer = bytearray()
async for chunk in audio_source:
if stop_event and stop_event.is_set(): break
buffer.extend(chunk)
while len(buffer) >= FRAME_SIZE:
await stream.send_audio(bytes(buffer[:FRAME_SIZE]))
buffer = buffer[FRAME_SIZE:]
await asyncio.sleep(frame_interval) # enforce real-time cadence
# Flush remaining partial frame padded with silence
if buffer:
await stream.send_audio(bytes(buffer) + b"\x00" * (FRAME_SIZE - len(buffer)))
```
***
## WebSocket Event Reference
| Event type | When sent | Key fields |
| ------------ | ----------------------------------------------- | ---------------------------------------------------------------------------------- |
| `connected` | Once, immediately after handshake. | `message`, `config.sample_rate`, `config.chunk_size`, `timestamp` |
| `processing` | Each time VAD detects end-of-speech. | `timestamp` |
| `transcript` | After transcription of a VAD segment completes. | `text`, `segment_index`, `segment_id`, `audio_duration_ms`, `latency`, `timestamp` |
| `error` | Server-side error, recoverable or fatal. | `message`, `timestamp` |
```json transcript event theme={null}
{
"type": "transcript",
"timestamp": "2024-01-15T10:30:05.987Z",
"text": "guaranteed returns milenge, bilkul risk-free hai",
"audio_duration_ms": 2340,
"segment_id": "seg_7f3a92",
"segment_index": 4,
"latency": 318
}
```
The `latency` field (milliseconds from end of speech to transcript delivery) is your primary observability metric for pipeline health. Track p50, p95, p99 per call session and alert if p95 consistently exceeds your SLA threshold.
***
## Compliance Detection
The compliance engine runs on each `transcript` event. It checks segment text against three rule categories: prohibited keywords, risk phrases, and required disclosures. All checks are synchronous string operations — they complete in under 1ms per segment.
```json rules/compliance.json theme={null}
{
"prohibited_keywords": [
{
"rule_id": "PROH_001", "severity": "CRITICAL",
"keywords": ["guaranteed returns", "guaranteed profit", "no risk", "risk-free"],
"description": "SEBI-prohibited investment language"
},
{
"rule_id": "PROH_002", "severity": "CRITICAL",
"keywords": ["personal account", "off the books", "my account"],
"description": "Agent directing customer to off-channel transaction"
}
],
"risk_phrases": [
{
"rule_id": "RISK_001", "severity": "WARNING",
"phrases": ["cancel my policy", "close my account", "policy cancel"],
"description": "Churn risk signal"
},
{
"rule_id": "RISK_002", "severity": "WARNING",
"phrases": ["legal action", "consumer forum", "RBI complaint", "complaint"],
"description": "Regulatory complaint intent"
}
],
"required_disclosures": [
{
"rule_id": "DISC_001", "severity": "CRITICAL",
"must_contain_one_of": ["this call is being recorded", "call recording", "recorded for quality"],
"check_within_segments": 3,
"description": "Recording disclosure required within first 3 segments"
}
]
}
```
```python ComplianceEngine theme={null}
import json
from pathlib import Path
from typing import List, Dict
class ComplianceEngine:
def __init__(self, rules_path="rules/compliance.json"):
rules = json.loads(Path(rules_path).read_text())
self.prohibited = rules.get("prohibited_keywords", [])
self.risk_phrases = rules.get("risk_phrases", [])
self.disclosures = rules.get("required_disclosures", [])
self._disclosed = set()
def check(self, segment) -> List[Dict]:
text = segment.text.lower()
hits = []
for rule in self.prohibited:
for kw in rule["keywords"]:
if kw in text:
hits.append({"rule_id": rule["rule_id"], "severity": rule["severity"],
"matched": kw, "description": rule["description"],
"segment_idx": segment.segment_index, "text": segment.text})
break
for rule in self.risk_phrases:
for phrase in rule["phrases"]:
if phrase in text:
hits.append({"rule_id": rule["rule_id"], "severity": rule["severity"],
"matched": phrase, "description": rule["description"],
"segment_idx": segment.segment_index, "text": segment.text})
break
for rule in self.disclosures:
rid = rule["rule_id"]
if rid in self._disclosed: continue
if any(p in text for p in rule["must_contain_one_of"]):
self._disclosed.add(rid)
elif segment.segment_index >= rule["check_within_segments"]:
hits.append({"rule_id": rid, "severity": rule["severity"],
"matched": "MISSING_DISCLOSURE", "description": rule["description"],
"segment_idx": segment.segment_index, "text": ""})
self._disclosed.add(rid)
return hits
```
***
## Quality Monitoring
```json rules/quality.json theme={null}
{
"silence": { "threshold_seconds": 8 },
"escalation_phrases": [
"transfer to supervisor", "let me escalate", "i will get my supervisor"
],
"interruption": { "min_duration_ms": 300 },
"short_segment_ms": 500
}
```
```python QualityEngine theme={null}
import json
from datetime import datetime, timezone
from pathlib import Path
from typing import List, Dict, Optional
class QualityEngine:
def __init__(self, rules_path="rules/quality.json"):
rules = json.loads(Path(rules_path).read_text())
self.silence_threshold = rules["silence"]["threshold_seconds"]
self.escalation_phrases = [p.lower() for p in rules["escalation_phrases"]]
self.interruption_ms = rules["interruption"]["min_duration_ms"]
self.short_segment_ms = rules["short_segment_ms"]
self._last_processing_ts: Optional[datetime] = None
def on_processing(self, timestamp_str: str):
self._last_processing_ts = datetime.fromisoformat(timestamp_str.replace("Z", "+00:00"))
def check(self, session, segment) -> List[Dict]:
now, text, events = datetime.now(timezone.utc), segment.text.lower(), []
if session.last_segment_end:
silence_s = (now - session.last_segment_end).total_seconds() - (segment.audio_duration_ms / 1000)
if silence_s > self.silence_threshold:
events.append({"event_type": "SILENCE", "severity": "WARNING",
"silence_s": round(silence_s, 1), "segment_idx": segment.segment_index,
"description": f"Silence gap of {silence_s:.1f}s detected"})
for phrase in self.escalation_phrases:
if phrase in text:
events.append({"event_type": "ESCALATION", "severity": "WARNING",
"matched": phrase, "segment_idx": segment.segment_index,
"description": "Supervisor escalation signal"})
break
if self._last_processing_ts and segment.audio_duration_ms < self.short_segment_ms:
gap_ms = (now - self._last_processing_ts).total_seconds() * 1000
if gap_ms < self.interruption_ms:
events.append({"event_type": "INTERRUPTION", "severity": "INFO",
"gap_ms": round(gap_ms, 1), "segment_idx": segment.segment_index,
"description": f"Possible interruption — {gap_ms:.0f}ms gap"})
return events
```
***
## Error Handling & Reconnect Logic
WebSocket connections drop. The reconnect loop below uses exponential backoff with full jitter and caps at a configurable maximum. Session state is preserved across reconnects using `processed_indices` to deduplicate segments.
```python reconnect loop theme={null}
import asyncio, random, os
from gnani.stt import GnaniSTTStreamClient, StreamConnectionError, StreamClosedError, StreamError
MAX_RECONNECTS = 5
BASE_BACKOFF_S = 1.0
MAX_BACKOFF_S = 30.0
async def monitor_call_with_reconnect(session, audio_source, compliance_engine, quality_engine, alert_dispatcher):
attempt = 0
while attempt <= MAX_RECONNECTS:
try:
async with GnaniSTTStreamClient(
api_key=os.getenv("GNANI_API_KEY"),
language_code=session.language_code,
sample_rate=int(os.getenv("SAMPLE_RATE", "16000")),
) as stream:
if attempt > 0: session.reconnect_count += 1
attempt = 0 # reset backoff counter on successful connect
stop_event = asyncio.Event()
producer = asyncio.create_task(stream_audio_producer(stream, audio_source, stop_event=stop_event))
async for event in stream:
await handle_event(session, event, compliance_engine, quality_engine, alert_dispatcher)
stop_event.set()
await producer
return # clean exit
except StreamConnectionError:
print(f"[{session.call_id}] Auth failure. Not retrying.")
raise
except (StreamClosedError, ConnectionResetError, OSError) as e:
attempt += 1
if attempt > MAX_RECONNECTS: raise
backoff = min(BASE_BACKOFF_S * (2 ** attempt), MAX_BACKOFF_S)
jitter = random.uniform(0, backoff * 0.2)
print(f"[{session.call_id}] Reconnect {attempt}/{MAX_RECONNECTS} in {backoff+jitter:.1f}s")
await asyncio.sleep(backoff + jitter)
```
| Error | Cause | Strategy |
| -------------------------------- | --------------------------------------------------------- | -------------------------------------------------------------------- |
| `StreamConnectionError` | 401, invalid API key, unsupported language code. | Do not retry. Fix config and redeploy. |
| `StreamClosedError` | Server closed cleanly (service restart, session timeout). | Retry with backoff. Session state is preserved. |
| `ConnectionResetError / OSError` | Network drop, TCP reset, intermediary timeout. | Exponential backoff + jitter. Cap at `MAX_RECONNECTS`. |
| `StreamError` | STT engine failure reported in an `error` event. | Log, retry once. Flag the call for manual review on repeat failures. |
***
## Production Best Practices
Each active call runs in its own `asyncio.Task`. The audio producer and event consumer run concurrently within that task. Do not use threads — the WebSocket library is async-native. A single well-tuned Python process handles 100+ concurrent calls comfortably; the bottleneck is network I/O, not CPU.
Compliance and quality checks run synchronously (sub-millisecond string matching). Alert dispatch — HTTP webhooks, queue publishes, database writes — must always be fire-and-forget via `asyncio.create_task()`. A slow downstream system under load must never delay the next transcript event.
| Optimization | Impact |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| Co-locate with telephony bridge | Run the monitor in the same region as the Gnani Prisma v2.5 API. Cross-region adds 50–150ms RTT per frame delivery. |
| 16kHz over 8kHz when possible | Higher accuracy transcripts mean fewer false positives in compliance matching. |
| Pre-compile compliance patterns | Compile all regex at engine `__init__`. Never compile inside the hot path. |
| Buffer writes, not reads | Write to an in-memory session buffer. Flush to the database at call end or on CRITICAL alerts only. |
| Metric | Source |
| ----------------------- | -------------------------------------------------------------- |
| `transcript_latency_ms` | `latency` field on each `transcript` event. Track p50/p95/p99. |
| `segment_count` | Increment on each `transcript` event. |
| `compliance_hit_rate` | Compliance hits / total segments per call. |
| `silence_gap_seconds` | Max silence gap derived from `processing` event timestamps. |
| `reconnect_count` | `session.reconnect_count`, incremented on each reconnect. |
***
## Debugging
| Symptom | Cause | Fix |
| ---------------------------------------------------- | ----------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
| Connection immediately closes — no `connected` event | Invalid API key, wrong `lang_code`, missing required headers. | Log the WebSocket close code — 4001 = auth failure. |
| Transcripts arrive but text is empty or garbled | `x-sample-rate` does not match the actual audio sample rate. Audio is not mono PCM. | Run `ffprobe` on the source. Convert stereo to mono before streaming. |
| VAD fires too often — sentences cut mid-utterance | Frames being burst-sent faster than real time. | Enforce `asyncio.sleep(frame_interval)` after every send. |
| VAD never fires — no processing or transcript events | Audio buffer is all zeros. Audio source is not connected. | Print `frame[:32].hex()`. All zeros = silent source. |
| Compliance rules fire on unrelated text | Substring match without word boundaries. | Switch to word-boundary regex. Lowercase and strip punctuation before matching. |
| Duplicate alerts on reconnect | Session buffer re-processed after reconnect. | Check `segment_index in session.processed_indices` before dispatching any alert. |
***
## Full Runnable Example
```python monitor.py theme={null}
"""
monitor.py — Real-time quality and compliance monitoring pipeline.
Usage:
GNANI_API_KEY=your-key python monitor.py --audio call.pcm --lang hi-IN
GNANI_API_KEY=your-key python monitor.py --audio call.pcm --lang en-IN --rate 8000
Install:
pip install gnani-vachana aiohttp python-dotenv
"""
import asyncio, json, os, random, argparse, aiohttp
from dataclasses import dataclass, field
from datetime import datetime, timezone
from pathlib import Path
from typing import List, Dict, Optional, Set, AsyncIterator
from dotenv import load_dotenv
from gnani.stt import (
GnaniSTTStreamClient,
StreamConnectedEvent, StreamProcessingEvent,
StreamTranscriptEvent, StreamErrorEvent,
StreamConnectionError, StreamClosedError, StreamError,
)
load_dotenv()
FRAME_SIZE = 1024
MAX_RECONNECTS = 5
BASE_BACKOFF_S = 1.0
MAX_BACKOFF_S = 30.0
FRAME_MS_16K = 0.032
FRAME_MS_8K = 0.064
@dataclass
class TranscriptSegment:
segment_index: int
text: str
audio_duration_ms: int
latency_ms: int
timestamp: datetime
compliance_flags: List[str] = field(default_factory=list)
quality_flags: List[str] = field(default_factory=list)
@dataclass
class CallSession:
call_id: str
language_code: str
started_at: datetime = field(default_factory=lambda: datetime.now(timezone.utc))
segments: List[TranscriptSegment] = field(default_factory=list)
last_segment_end: Optional[datetime] = None
reconnect_count: int = 0
processed_indices: Set[int] = field(default_factory=set)
def add_segment(self, event) -> TranscriptSegment:
seg = TranscriptSegment(
segment_index=event.segment_index, text=event.text,
audio_duration_ms=event.audio_duration_ms, latency_ms=event.latency,
timestamp=datetime.now(timezone.utc),
)
self.segments.append(seg)
self.processed_indices.add(event.segment_index)
self.last_segment_end = datetime.now(timezone.utc)
return seg
async def stream_audio_producer(stream, audio_source, stop_event: asyncio.Event):
buffer = bytearray()
async for chunk in audio_source:
if stop_event.is_set(): break
buffer.extend(chunk)
while len(buffer) >= FRAME_SIZE:
await stream.send_audio(bytes(buffer[:FRAME_SIZE]))
buffer = buffer[FRAME_SIZE:]
if buffer:
await stream.send_audio(bytes(buffer) + b" " * (FRAME_SIZE - len(buffer)))
async def handle_event(session, event, compliance_engine, quality_engine, alert_dispatcher):
if isinstance(event, StreamConnectedEvent):
print(f"[{session.call_id}] Connected sample_rate={event.sample_rate}")
elif isinstance(event, StreamProcessingEvent):
quality_engine.on_processing(event.timestamp)
elif isinstance(event, StreamTranscriptEvent):
if event.segment_index in session.processed_indices:
return # deduplicate across reconnects
seg = session.add_segment(event)
print(f"[{session.call_id}][{seg.segment_index}] {seg.text} (latency={seg.latency_ms}ms)")
c_hits = compliance_engine.check(seg)
q_events = quality_engine.check(session, seg)
seg.compliance_flags = [h["rule_id"] for h in c_hits]
seg.quality_flags = [e["event_type"] for e in q_events]
if c_hits or q_events:
asyncio.create_task(alert_dispatcher.send(session, seg, c_hits + q_events))
elif isinstance(event, StreamErrorEvent):
raise RuntimeError(f"STT error: {event.message}")
async def monitor_call(session, audio_source, compliance_engine, quality_engine, alert_dispatcher):
attempt = 0
while attempt <= MAX_RECONNECTS:
try:
async with GnaniSTTStreamClient(
api_key=os.getenv("GNANI_API_KEY"),
language_code=session.language_code,
sample_rate=int(os.getenv("SAMPLE_RATE", "16000")),
) as stream:
if attempt > 0: session.reconnect_count += 1
attempt = 0
stop_event = asyncio.Event()
producer = asyncio.create_task(stream_audio_producer(stream, audio_source, stop_event))
async for event in stream:
await handle_event(session, event, compliance_engine, quality_engine, alert_dispatcher)
stop_event.set()
await producer
return
except StreamConnectionError: raise
except (StreamClosedError, ConnectionResetError, OSError) as e:
attempt += 1
if attempt > MAX_RECONNECTS: raise
backoff = min(BASE_BACKOFF_S * (2 ** attempt), MAX_BACKOFF_S)
print(f"[{session.call_id}] Reconnect {attempt}/{MAX_RECONNECTS} in {backoff:.1f}s {e}")
await asyncio.sleep(backoff + random.uniform(0, backoff * 0.2))
except StreamError as e:
attempt += 1
print(f"[{session.call_id}] Server error: {e.message}")
await asyncio.sleep(BASE_BACKOFF_S)
async def file_audio_source(path: str, sample_rate=16000) -> AsyncIterator[bytes]:
frame_interval = FRAME_MS_16K if sample_rate == 16000 else FRAME_MS_8K
with open(path, "rb") as f:
while chunk := f.read(FRAME_SIZE):
yield chunk
await asyncio.sleep(frame_interval)
async def main():
parser = argparse.ArgumentParser()
parser.add_argument("--audio", required=True)
parser.add_argument("--lang", default="hi-IN")
parser.add_argument("--rate", default=16000, type=int)
args = parser.parse_args()
os.environ["SAMPLE_RATE"] = str(args.rate)
session = CallSession(call_id="CALL_001", language_code=args.lang)
audio = file_audio_source(args.audio, sample_rate=args.rate)
await monitor_call(session, audio, ComplianceEngine(), QualityEngine(), AlertDispatcher())
duration = (datetime.now(timezone.utc) - session.started_at).total_seconds()
print(f"Duration: {duration:.1f}s | Segments: {len(session.segments)}")
print(f"Compliance hits: {sum(len(s.compliance_flags) for s in session.segments)}")
print(f"Quality events: {sum(len(s.quality_flags) for s in session.segments)}")
if __name__ == "__main__":
asyncio.run(main())
```
***
## What to Build Next
* **Speaker Diarization** — Separate agent and customer voices. Attribute compliance hits to the correct speaker.
* **Sentiment Analysis** — Feed each segment's text to a sentiment model. Track the sentiment arc across the call.
* **Agent Assist** — On each `transcript` event, call an LLM with the running conversation context to surface next-best-action suggestions in real time.
* **LLM Summarisation** — At call end, send the full session transcript to an LLM for structured output: issue, resolution, action items, disposition.
* **Compliance Scoring** — Build a per-call compliance score (0–100) based on rule severity, frequency, and placement in the call.
**Related docs:** [WebSocket STT API](/vachana/STT/stt-websocket) · [Batch STT for post-call analysis](/vachana/STT/stt-batch) · SDK install: `pip install gnani-vachana`
# Calls
Source: https://docs.gnani.ai/b
The **Calls** section provides a centralised view of all call recordings available for analysis and evaluation. Each call is processed to extract key insights, including sentiment, emotion, language, and quality scores.
From this section, users can:
* View processed calls
* Upload call recordings
* Automatically ingest calls through integrations
* Review analytics and evaluation results
This page acts as the starting point for reviewing and evaluating agent interactions.
***
# Viewing Calls
## Overview
The **Viewing Calls** page displays a list of all call recordings that have been uploaded or ingested into the system.
Each row in the table represents a call and includes important information related to the interaction.
***
## Call List
The call list provides a structured overview of all interactions along with key metadata and analytics.
Information displayed in the call list includes:
* **Call ID** – Unique identifier for the call
* **Recording Name** – Name of the uploaded audio file
* **Processed Date**
* **Agents** - Call performed by a certain agent
* **Status** – Processing status of the call\
**Key Metrics** – Important tags or indicators detected during processing
* **Call Duration** – Length of the conversation
* **Language** – Language detected or selected for the call
* **Customer Sentiment** – Sentiment detected from the customer
* **Agent Sentiment** – Sentiment detected from the agent
* **Emotion** – Emotion detected in the conversation
* **QA Call Score** – Score assigned during evaluation
* **Audio Recording**
* **Actions**
* **View Scorecard CTA**
## Call Actions
Users can perform actions directly from the call list.
These actions include:
**Play Recording-** Play the call recording directly from the table to quickly review the interaction.
**Retry Processing -** If a call fails to process correctly, users can retry processing.
**View Evaluation -** Open the scorecard evaluation associated with the call.
***
# Uploading Calls
## Overview
Call recordings can be uploaded manually for processing and analysis.
Once uploaded, the system processes the audio to generate insights and make the call available for evaluation.
## Steps to Upload Calls
1. Navigate to the **Calls** page
2. Click **Upload Calls**
3. Upload the call recording files
Supported audio formats include:
* MP3
* WAV
* AAC
* M4A
## Assigning an Agent
During upload, each call must be associated with an agent.
Steps:
1. Select **Assign an Agent**
2. Choose the agent responsible for the call
This ensures the call and evaluation results are linked to the correct agent.
***
## Selecting the Language
Users must select the language of the call recording.
This allows accurate processing and analysis of the audio.
***
## Uploading Audio Files
Users can upload files by either:
* Clicking **Upload** and selecting files from the device
* Dragging and dropping files into the upload area
Once uploaded, the call is sent for processing and appears in the call list.
# Call Analysis
Source: https://docs.gnani.ai/c
## Overview
Once call recordings are uploaded and processed, detailed insights become available in the **Call Analysis** view. This page provides a comprehensive breakdown of the interaction, including the conversation transcript, sentiment analysis, speech metrics, and AI-generated summaries.
The analysis helps reviewers quickly understand the context of the conversation, identify key events during the call, and evaluate agent performance more efficiently.
***
# AI Analytics
The **AI Analytics** tab provides automatically generated insights about the call based on the conversation between the agent and the customer.
These insights help users understand the purpose of the call, how it progressed, and how effectively the interaction was handled.
***
## 1. Call Summary
The **Call Summary** section highlights the most important aspects of the conversation.
It includes the following insights:
### 1.1 Reason of the Call
This describes the primary purpose of the interaction based on the conversation between the agent and the customer.
Example: The primary reason for the interaction is to initiate a conversation regarding financial services with the customer.
***
### 1.2 Resolution
This indicates whether the agent addressed the customer's request and how the conversation concluded.\
It summarises the outcome of the interaction.
***
### 1.3 Summary
This provides a short narrative describing how the conversation progressed from start to end.
It typically includes:
* Greeting and introduction
* Discussion between the participants
* Key conversation points
* Outcome of the interaction
This helps reviewers quickly understand the call without listening to the full recording.
***
### 1.4 Feedback
The feedback section highlights possible improvements in the interaction.
These insights may include suggestions such as:
* Improving communication clarity
* Providing additional context
* Enhancing customer engagement
This helps supervisors identify coaching opportunities.
***
### 1.5 Call Forwarding
This indicates whether the call was transferred or forwarded during the interaction.
If the call was not transferred, the value is displayed as **No**.
***
## 2. Recording & Transcript
The **Recording & Transcript** section allows users to review the conversation in detail.
***
## 3. Call Recording
The audio player allows users to listen to the call recording directly within the interface.
Features include:
* Play and pause controls
* Playback timeline
* Adjustable playback speed
This allows evaluators to review the call while analysing the conversation.
***
## 4. Transcript
The transcript displays the conversation between the agent and the customer in a chat-style format.
Each message is tagged with:
* **Speaker identification** (Agent or Customer)
* **Sentiment** detected for that message
* **Emotion** associated with the response
This helps reviewers understand the tone and emotional context of the conversation.
***
## 5. Sentiment Indicators
The transcript also highlights sentiment detected in each message.
Examples include:
* Positive
* Neutral
* Negative
This helps identify moments in the conversation where the customer's sentiment may have changed.
***
## 6. Rate of Speech
The **Rate of Speech** metric measures how fast the agent is speaking during the conversation.
This is typically measured in **words per minute (WPM)**.
Monitoring speech rate helps identify communication issues such as:
* Speaking too quickly for customers to understand
* Long pauses during conversation
Maintaining an appropriate speaking pace improves customer experience.
***
## 7. Key Indicators
Key Indicators allow organizations to track specific conversation signals that are important for quality monitoring.
These indicators may represent:
* Compliance phrases
* Required statements
* Product mentions
* Process confirmations
Indicators can be configured to detect when these events occur during the call.
***
## 8. Key Metrics
The **Key Metrics** section highlights important operational metrics related to the interaction.
These metrics help understand how effectively the conversation progressed.
Examples of metrics may include:
* Conversation duration
* Key interaction signals
* Process adherence indicators
***
## 9. Advanced Metrics
The **Advanced Metrics** section provides deeper analytics about the interaction.
These metrics help analyze conversation flow and communication patterns.
Metrics include:
### 9.1 Initial Language
The language detected at the beginning of the conversation.
***
### 9.2 Number of Turns
The number of conversational exchanges between the agent and the customer.
Each turn represents a switch between speakers.
***
### 9.3 Number of Nudges
This metric indicates the number of times the system detected potential prompts or cues that could guide the conversation.
***
### 9.4 Dead Air
Dead air represents periods of silence during the conversation where no one is speaking.
Long periods of silence may indicate issues such as:
* Delays in response
* System interruptions
* Agent hesitation
***
### 9.5 Interest Conversion Turns
This metric measures the number of conversation turns where customer interest or engagement increased during the call.
***
### 9.6 Average Customer Response Time
The average time taken by the customer to respond during the conversation.
***
### 9.7 Average Agent Response Time
The average time taken by the agent to respond to the customer.
This helps evaluate responsiveness during the interaction.
***
## 10. Key Topics
The **Key Topics** section highlights frequently used words or topics that appeared in the conversation.
These topics are presented as a visual word cluster to help reviewers quickly identify the dominant themes of the interaction.
This provides a quick overview of the discussion topics without reviewing the full transcript.
# QA Analytics
## Overview
The **QA Analytics** section provides a structured evaluation of the call using a scorecard form. This evaluation measures the quality of the interaction based on predefined criteria such as greeting, communication clarity, professionalism, and process adherence.
When a call is processed, the system automatically applies the **default scorecard form** and generates an evaluation score based on the responses detected from the conversation.
This allows reviewers to quickly assess how well the agent adhered to the expected call-handling guidelines.
***
## Scorecard Evaluation
At the top of the page, the **scorecard used for evaluation** is displayed.
The scorecard contains multiple sections, each representing a category of evaluation.
Examples of sections include:
* Call Opening
* Soft Skills & Professionalism
* Communication Skills
Each section contains a set of questions used to assess specific aspects of the interaction.
***
## AI Score
The **AI Score** represents the overall quality score generated from the scorecard evaluation.
This score is calculated based on the marks assigned to each question in the scorecard.
Example:
AI Score: **35 / 100**
This score provides a quick indicator of how well the agent performed during the interaction.
***
## Scorecard Questions
Each section contains multiple questions that evaluate specific behaviors during the call.
Examples include:
* Did the agent greet the customer and introduce themselves appropriately?
* Did the agent clearly state the purpose of the call?
* Was the agent's tone friendly and professional?
* Did the agent inform the customer that the call is being recorded?
For each question, the system assigns a response based on the conversation analysis.
Possible responses include:
* **Yes**
* **No**
* **N/A (Not Applicable)**
Each response corresponds to a specific mark value.
***
## Question Scoring
Every question in the scorecard has predefined marks.
Example scoring:
* Yes → Full marks
* No → Zero marks
* N/A → Question excluded from scoring
The marks assigned to each response contribute to the **total section score**.
***
## Section Scores
Each scorecard section has a **maximum mark value**.
Example: Call Opening – Total Marks: **10**
The marks obtained from questions within that section are aggregated to generate the **section score**.
Example: 5 / 10 marks
This helps reviewers understand which parts of the interaction were handled well and which areas need improvement.
***
## Collapsible Sections
Scorecard sections are collapsible, allowing users to expand or collapse sections as needed.
This helps reviewers focus on specific parts of the evaluation without scrolling through the entire scorecard.
***
# Using QA Analytics
The QA Analytics view helps organizations:
* Monitor adherence to call handling guidelines
* Evaluate agent performance consistently
* Identify training opportunities
* Improve overall customer interaction quality
By combining scorecard evaluations with call analytics, reviewers can understand both **what happened during the call and how well it was handle**
# Creating a Workforce
Source: https://docs.gnani.ai/creating-a-workforce
## What is Workforce?
Workforce allows you to create a **multi-agent workflow** where multiple AI voice agents work together to handle different parts of a conversation.
Instead of building one large agent that handles every possible scenario, Workforce lets you connect multiple specialized agents together. Each agent focuses on a specific task, and conversations move between them based on defined conditions.
For example:
* A Greeting Agent that welcomes the user and gathers basic information
* A Support Agent that answers product questions
* A Cancellation Agent that handles refund or cancellation requests
By distributing responsibilities across agents, workflows become easier to maintain, more scalable and more reliable.
Think of Workforce as building a team of AI employees working together in a coordinated flow.
***
## Why Use Workforce?
A single agent managing too many tasks can become:
* Hard to maintain
* Difficult to debug
* Less reliable in complex scenarios
Workforce solves this by enabling specialized agents connected through logic-based routing.
Benefits include:
1. Modular Design: Each agent handles a specific responsibility.
2. Better Maintainability: Update one agent without affecting others.
3. Scalable Workflows: Add new agents as your use case grows.
4. Smarter Routing: Conversations move automatically between agents based on conditions.
***
## Key Concepts
### Workforce Canvas
The Workforce Canvas is the visual workspace where you design your multi-agent flow.
On the canvas you can:
* Add agents
* Connect agents together
* Define conditions for transferring conversations
Each element on the canvas is represented as a node.
### Agent Nodes
Agent nodes represent the individual AI agents that handle conversations.
Each agent performs a specific role in the workflow.
Agents added to the workforce are imported from your existing agent library.
### Edges (Connections)
Edges connect agents together and define how conversations move from one agent to another.
Each connection can include a condition that determines when the handoff should occur.
### Trigger Node
The Trigger Node is the starting point of the workforce.
As of now, every workforce begins with an Outbound Call Trigger that is automatically added to the canvas when the workforce is created.
This trigger represents the starting point of the workflow. From this node, you can connect the first agent that should handle the conversation.
***
## A. Creating a Workforce
1. Navigate to **Build ->** **Workforce** section in the platform.
2. Click **+ Workforce** button.
3. You will be asked to provide the name and description.
4. After entering the details, click **Create**.
5. You will be redirected to the **Workforce Canvas**.
A trigger node will already be present on the canvas.
***
## B. Adding Agents to the Workforce
Agents used inside a workforce are imported from your existing agent library.
### How to Import an Agent
1. On the Workforce Canvas, select the option to add an agent.
2. Select an agent from your agent library.
3. Confirm the selection.
The selected agent will appear as a node on the canvas.
### Linked Agent Behavior
Agents imported into a workforce are **linked to the original agent in the Manage Agents library**.
This means:
* The agent configuration cannot be edited from the workforce canvas.
* Changes made to the agent in the Manage Agents page will automatically apply to all workforces where the agent is used.
### Editing an Imported Agent
To modify the agent configuration:
1. Navigate to Manage Agents.
2. Open the agent.
3. Make the required changes.
4. Save the agent.
The updated behavior will automatically reflect in all workforces where the agent is imported.
***
## C. Connecting Agents
Agents must be connected together to define how conversations move through the workflow.
### Creating a Connection
1. Drag a connector from Agent A to Agent B.
2. A connection will be created between the two agents.
After the connection is created, you can define the condition that determines when the handoff should occur.
***
## D. Conditional Handoff Between Agents
Conditional handoff allows conversations to move between agents when specific conditions are met.
These conditions are written in natural language.
### How to Configure a Conditional Handoff
1. Drag a connector from Agent A to Agent B.
2. A label saying Add condition appears on the connection.
3. Click the label.
4. A configuration sidebar opens.
5. Enter the condition and edge label.
6. Save the configuration.
### Example Conditions
Examples of natural language conditions include:
* "User asks for a supervisor"
* "User wants to cancel the appointment"
* "Sentiment is negative and the user mentions cancel"
### How the Handoff Works During a Call
If the condition defined on the connection is met during the conversation:
* The conversation will automatically transfer to the next agent.
If the condition is never triggered:
* The transfer will not occur.
### Data Passed During Handoff
When a conversation is transferred between agents, the system passes the complete context to the next agent, including:
* Conversation transcript
* Structured data extracted during the call
* Session context
This allows the next agent to continue the conversation without losing information.
***
## Best Practices
### Use Specialized Agents
Each agent should handle a specific responsibility. Avoid building agents that handle **too many tasks**.
Example:
* Greeting Agent
* Support Agent
* Cancellation Agent
This keeps workflows simpler and easier to maintain.
### Write Clear Handoff Conditions
Use clear and specific conditions when defining agent transitions.
Examples:
* "User wants to cancel subscription"
* "User requests appointment rescheduling"
* "User asks to speak to supervisor"
Avoid vague conditions.
### Reuse Agents Across Workforces
If multiple workflows require the same agent behavior, reuse the same agent by importing it into different workforces.
Any improvements made to that agent will automatically apply everywhere it is used.
***
## Summary
The Workforce feature allows you to design multi-agent workflows where specialized agents collaborate to handle conversations.
With Workforce you can:
* Create workflows composed of multiple agents
* Visually design conversation flows
* Transfer conversations between agents based on conditions
* Reuse agents across multiple workflows
This modular approach makes complex voice automation systems easier to build, manage, and scale.
# Introduction
Source: https://docs.gnani.ai/introduction
Welcome, future innovator! This guide is your friendly roadmap to building a powerful AI agent, even if you're totally new to conversational AI. Here, you'll learn not only how to set up your agent but also why each feature is important and how it can be applied to real-world scenarios. Whether you're automating customer support, generating leads, creating a smart scheduling assistant, or simply curious about AI, we’re here to make it fun and simple 🚀
## Setting up
### *Welcome Aboard!*
You’re about to build a conversational GenAI agent that can talk with your customers, understand, empathize and provide solutions. Think of this platform as your agentic lab.
* **Knowledge Bases** = Your agent’s memory.
* **Agents** = Your 24/7 delegates who speak with your customers.
* **Integrations** = Superpowers (like SMS, email, CRM).
* **Actions** = Your agent’s workflows and automations.
**First Steps:**
1. **Create a Knowledge Base** (Feed your agent information).
2. **Build an Agent** (Define its role and style).
3. **Add Integrations** (Connect to tools like Zoho or Twilio).
4. **Add Actions** (Define how your agent uses integrations to automate tasks).
5. **Test & Launch** (Watch it come alive!).
The first step to world-class documentation is setting up your editing environments.
Learn how to upload documents or import content from a URL to build your AI agent's knowledge base
Learn how to create and customize your first GenAI agent, from scratch or using a template
## Personalize Your Agent
Tailor your GenAI agent to meet your specific needs by customizing
Learn how to integrate your GenAI agent with SMS, email, CRM, ticketing systems, and custom APIs
Learn how to create variables and actions to automate workflows and handle dynamic data in your GenAI agent
Create multi-step conversational flows that adapt to user needs with Agent Chaining
Learn how to whitelist phone numbers for testing and import numbers from Twilio to your agent
Learn how to track and improve your agent’s performance with conversational and action logs, as well as in-depth analytics
Easily raise support tickets for bugs, features, or account issues, and track the progress of your requests
# Introduction
Source: https://docs.gnani.ai/introduction-1
Welcome to **Aura**, a conversation analytics and quality monitoring platform designed to help teams understand, evaluate, and improve customer interactions.
Aura automatically analyses customer calls to generate actionable insights about conversations, agent performance, and customer sentiment. Instead of manually reviewing calls, teams can rely on AI-powered analysis to quickly identify patterns, detect important signals, and monitor quality at scale.
Whether you are reviewing individual calls or analysing thousands of conversations, Aura helps you transform raw call data into meaningful insights.
# What You Can Do With Aura
Aura provides a comprehensive set of tools for analysing and evaluating customer conversations.
With Aura, you can:
* **Analyse customer calls** using AI-generated transcripts, summaries, and insights
* **Evaluate agent performance** using customizable scorecard forms
* **Track important conversation signals** using key indicators and metrics
* **Monitor call quality** and detect potential issues such as silence or interruptions
* **Extract insights from conversations** using AI-powered analytics
* **Identify trends across large volumes of calls**
These capabilities help teams improve customer experience, agent training, and operational performance.
## First Steps
Getting started with Aura is simple. Follow these steps to begin analysing your conversations.
1. **Add Users**\
Invite agents, QA analysts, or supervisors to the platform.
2. **Create Scorecard Forms**\
Build evaluation templates to measure agent performance and call quality.
3. **Configure Analytics Settings**\
Define key indicators, metrics, and AI analytics to extract insights from conversations.
4. **Upload or Connect Calls**\
Upload call recordings manually or connect your call storage through integrations.
5. **Review Call Insights**\
Explore transcripts, summaries, and conversation analytics generated by the system.
6. **Evaluate Calls**\
Use scorecards to review calls and monitor performance across teams.
***
## What Happens Next
Once calls are uploaded, Aura automatically processes them and generates insights such as:
* call summaries
* transcripts
* sentiment analysis
* detected keywords and phrases
* conversation topics
* performance metrics
These insights appear across the **Dashboard and Call Analytics**, helping your team quickly understand what happened in every conversation.
# LiveKit Plugin
Source: https://docs.gnani.ai/livekit/introduction
Use Gnani Prisma v2.5 (STT) and Gnani Timbre v2.0 (TTS) inside LiveKit Agents voice pipelines for Indian languages.
## Overview
`livekit-plugins-gnani` is a thin LiveKit Agents adapter that wraps the Gnani STT and TTS APIs into LiveKit's standard `stt.STT` and `tts.TTS` base classes. Drop it into any LiveKit voice agent pipeline and get high-accuracy Indian-language speech recognition and low-latency synthesis without managing WebSocket connections yourself.
```text theme={null}
gnani-vachana ← Core SDK (REST, WebSocket, SSE clients)
↑
livekit-plugins-gnani ← This package (LiveKit Agents adapter)
↑
Your LiveKit voice agent
```
All connection logic, authentication, and audio format handling live in the core SDK. The plugin is purely an adapter layer.
***
## Installation
```bash theme={null}
pip install livekit-plugins-gnani
```
This also installs `gnani-vachana` (the core SDK) as a dependency.
**Requirements:** Python 3.10+
***
## Prerequisites
You need a Gnani API key. [Gnani APIs](https://app.gnani.ai/voice)
Set your credentials as environment variables:
```bash theme={null}
export GNANI_API_KEY="your-api-key"
# For REST STT only (optional — streaming STT and TTS need only GNANI_API_KEY):
export GNANI_ORGANIZATION_ID="your-org-id"
export GNANI_USER_ID="your-user-id"
```
***
## Quick Start
### Speech-to-Text: Gnani Prisma v2.5
```python theme={null}
from livekit.plugins.gnani import STT
stt = STT(language="hi-IN")
# Pass to your LiveKit voice agent pipeline
```
### Text-to-Speech: Gnani Timbre v2.0
```python theme={null}
from livekit.plugins.gnani import TTS
tts = TTS(voice="sia")
# Pass to your LiveKit voice agent pipeline
```
***
## STT — Speech-to-Text (Gnani Prisma v2.5)
The plugin exposes two STT modes matching the underlying Gnani Prisma v2.5 API:
| Mode | Method | Best for |
| --------------------- | ------------------ | ---------------------------------------- |
| Streaming (WebSocket) | Default | Live conversations, real-time agents |
| Batch (REST) | `STT(mode="rest")` | Pre-recorded audio, file-based pipelines |
### Streaming (default)
Real-time transcription via WebSocket with Voice Activity Detection. This is the recommended mode for live voice agents.
```python theme={null}
from livekit.plugins.gnani import STT
stt = STT(
language="hi-IN", # BCP-47 language code
sample_rate=16000, # 8000 or 16000 Hz
)
```
### Batch (REST)
File-based transcription in a single synchronous request. Audio clips up to 60 seconds.
```python theme={null}
from livekit.plugins.gnani import STT
stt = STT(
language="hi-IN",
mode="rest",
)
```
### Code-switching (multilingual)
For audio that mixes Hindi and English, use the experimental Hinglish codes:
```python theme={null}
# Code-mixed: Latin + Devanagari output
stt = STT(language="en-hi-in-cm")
# Hinglish in Latin script only
stt = STT(language="en-hi-IN-latn")
```
***
## TTS — Text-to-Speech (Gnani Timbre v2.0)
The plugin exposes Gnani Timbre v2.0 TTS synthesis with 8 Indian voices.
```python theme={null}
from livekit.plugins.gnani import TTS
tts = TTS(
voice="sia", # Voice ID — see Available Voices below
sample_rate=16000, # Output sample rate: 8000–44100 Hz
encoding="linear_pcm", # linear_pcm or oggopus
)
```
### Available Voices
| Voice | ID |
| ------ | -------- |
| Sia | `sia` |
| Raju | `raju` |
| Kanika | `kanika` |
| Nikita | `nikita` |
| Ravan | `ravan` |
| Simran | `simran` |
| Karan | `karan` |
| Neha | `neha` |
***
## Supported Languages
| Language | Code |
| ------------------------------- | --------------- |
| Bengali | `bn-IN` |
| English (India) | `en-IN` |
| Gujarati | `gu-IN` |
| Hindi | `hi-IN` |
| Kannada | `kn-IN` |
| Malayalam | `ml-IN` |
| Marathi | `mr-IN` |
| Punjabi | `pa-IN` |
| Tamil | `ta-IN` |
| Telugu | `te-IN` |
| Hinglish *(experimental)* | `en-hi-in-cm` |
| Hinglish Latin *(experimental)* | `en-hi-IN-latn` |
***
## Further Reading
* [STT REST API](/vachana/STT/speech-to-text) — file-based transcription reference
* [STT Realtime API](/vachana/STT/stt-websocket) — WebSocket protocol reference
* [TTS REST API](/vachana/TTS/tts-inference) — synchronous synthesis reference
* [TTS Realtime API](/vachana/TTS/tts-websocket) — WebSocket TTS reference
* [`gnani-vachana` on PyPI](https://pypi.org/project/gnani-vachana/) — core SDK
* [`livekit-plugins-gnani` on PyPI](https://pypi.org/project/livekit-plugins-gnani/) — this plugin
* [LiveKit Agents Docs](https://docs.livekit.io/agents/) — LiveKit framework reference
# Pipecat Plugin
Source: https://docs.gnani.ai/pipecat/introduction
Use Gnani STT and TTS inside Pipecat voice agent pipelines for Indian languages.
## Overview
`pipecat-gnani` is a Pipecat service integration that wraps the Gnani STT and TTS APIs into Pipecat's standard `STTService`, `TTSService`, and `InterruptibleTTSService` base classes. Drop the services into any Pipecat pipeline and get high-accuracy Indian-language transcription and low-latency synthesis without managing WebSocket connections yourself.
```text theme={null}
gnani-vachana ← Core SDK (REST, WebSocket, SSE clients)
↑
pipecat-gnani ← This package (Pipecat service adapter)
↑
Your Pipecat voice agent
```
All connection logic, authentication, and audio format handling live in the core SDK. The plugin is purely an adapter layer.
***
## Installation
```bash theme={null}
pip install pipecat-gnani
```
This also installs `gnani-vachana` (the core SDK) as a dependency.
**Requirements:** Python 3.10+
***
## Prerequisites
You need a Gnani API key. [Gnani APIs](https://app.gnani.ai/voice)
```bash theme={null}
export GNANI_API_KEY="your-api-key"
```
***
## Services
This plugin provides three service classes. Choose based on your use case:
| Service | Type | Transport | Best for |
| --------------------- | ---- | ------------------- | ----------------------------------------------- |
| `GnaniSTTService` | STT | WebSocket | Live conversations, real-time agents |
| `GnaniTTSService` | TTS | WebSocket streaming | Conversational agents with interruption support |
| `GnaniHttpTTSService` | TTS | REST | Batch synthesis, non-streaming pipelines |
***
## Quick Start
### Speech-to-Text: Gnani Prisma v2.5
```python theme={null}
from pipecat_gnani import GnaniSTTService
from pipecat_gnani.language import Language
stt = GnaniSTTService(
api_key="your-api-key",
settings=GnaniSTTService.Settings(
language=Language.HI_IN,
),
)
```
### Text-to-Speech (WebSocket streaming — recommended): Gnani Timbre v2.0
```python theme={null}
from pipecat_gnani import GnaniTTSService
tts = GnaniTTSService(
api_key="your-api-key",
settings=GnaniTTSService.Settings(
voice="sia",
language="IND-IN",
),
)
```
### Text-to-Speech (REST): Gnani Timbre v2.0
```python theme={null}
from pipecat_gnani import GnaniHttpTTSService
tts = GnaniHttpTTSService(
api_key="your-api-key",
aiohttp_session=session, # pass your aiohttp.ClientSession here
settings=GnaniHttpTTSService.Settings(
voice="sia",
language="hi-IN",
),
)
```
***
## STT — `GnaniSTTService (Gnani Prisma v2.5)`
Real-time streaming speech-to-text via WebSocket with built-in Voice Activity Detection.
* Connects to `wss://api.vachana.ai/stt/v3/stream`
* Sends raw PCM audio in 1,024-byte frames
* Receives transcript events with segment metadata (`text`, `segment_id`, `segment_index`, `latency`)
* Supports 8 kHz and 16 kHz sample rates
```python theme={null}
from pipecat_gnani import GnaniSTTService
from pipecat_gnani.language import Language
stt = GnaniSTTService(
api_key="your-api-key",
settings=GnaniSTTService.Settings(
language=Language.HI_IN,
sample_rate=16000, # 8000 or 16000
),
)
```
### Settings
| Parameter | Type | Default | Description |
| ------------- | ---------- | ---------------- | --------------------------------------------------------------------------------- |
| `language` | `Language` | `Language.EN_IN` | Language enum for transcription. See [Supported Languages](#supported-languages). |
| `sample_rate` | `int` | `16000` | Audio sample rate in Hz. Accepted values: `8000`, `16000`. |
***
## TTS — `GnaniTTSService` (WebSocket, recommended): Gnani Timbre v2.0
Streaming text-to-speech via WebSocket. Extends Pipecat's `InterruptibleTTSService`, giving your agent built-in interruption (barge-in) support — when the user speaks over the agent, synthesis stops cleanly.
* Connects to `wss://api.vachana.ai/api/v1/tts`
* Streams audio chunks in real-time as synthesis progresses
* Ideal for live conversational agents where latency and barge-in handling matter
```python theme={null}
from pipecat_gnani import GnaniTTSService
tts = GnaniTTSService(
api_key="your-api-key",
settings=GnaniTTSService.Settings(
voice="sia",
language="IND-IN",
),
)
```
`GnaniTTSService` uses `"IND-IN"` as its language identifier, not the BCP-47 codes used elsewhere. This is a Gnani Timbre v2.0 WebSocket TTS protocol detail — the language selection is driven primarily by the `voice` parameter.
### Settings
| Parameter | Type | Default | Description |
| ------------- | -------- | ---------- | ---------------------------------------------------- |
| `voice` | `string` | `"sia"` | Voice ID. See [Available Voices](#available-voices). |
| `language` | `string` | `"IND-IN"` | Language identifier for the WebSocket TTS protocol. |
| `sample_rate` | `int` | `16000` | Output sample rate in Hz. |
***
## TTS — `GnaniHttpTTSService` (REST): Gnani Timbre v2.0
REST-based text-to-speech for non-streaming use cases. Returns the complete audio in a single response.
* Calls `POST /api/v1/tts/inference`
* Requires an active `aiohttp.ClientSession` passed at construction time
* Suitable for batch synthesis or pipelines where streaming is not needed
```python theme={null}
import aiohttp
from pipecat_gnani import GnaniHttpTTSService
async def build_pipeline():
async with aiohttp.ClientSession() as session:
tts = GnaniHttpTTSService(
api_key="your-api-key",
aiohttp_session=session,
settings=GnaniHttpTTSService.Settings(
voice="sia",
language="hi-IN",
),
)
```
### Settings
| Parameter | Type | Default | Description |
| ------------- | -------- | --------- | ---------------------------------------------------------------------- |
| `voice` | `string` | `"sia"` | Voice ID. See [Available Voices](#available-voices). |
| `language` | `string` | `"hi-IN"` | BCP-47 language code. See [Supported Languages](#supported-languages). |
| `sample_rate` | `int` | `16000` | Output sample rate in Hz. |
***
## Available Voices
| Voice | ID |
| ------ | -------- |
| Sia | `sia` |
| Raju | `raju` |
| Kanika | `kanika` |
| Nikita | `nikita` |
| Ravan | `ravan` |
| Simran | `simran` |
| Karan | `karan` |
| Neha | `neha` |
***
## Supported Languages
| Language | Code (`Language` enum) | BCP-47 string |
| --------------- | ---------------------- | ------------- |
| Bengali | `Language.BN_IN` | `bn-IN` |
| English (India) | `Language.EN_IN` | `en-IN` |
| Gujarati | `Language.GU_IN` | `gu-IN` |
| Hindi | `Language.HI_IN` | `hi-IN` |
| Kannada | `Language.KN_IN` | `kn-IN` |
| Malayalam | `Language.ML_IN` | `ml-IN` |
| Marathi | `Language.MR_IN` | `mr-IN` |
| Punjabi | `Language.PA_IN` | `pa-IN` |
| Tamil | `Language.TA_IN` | `ta-IN` |
| Telugu | `Language.TE_IN` | `te-IN` |
Use the `Language` enum for `GnaniSTTService`. Use the BCP-47 string for `GnaniHttpTTSService`. `GnaniTTSService` uses `"IND-IN"` regardless of language — voice selection handles language implicitly.
***
## Further Reading
* [STT REST API](/vachana/STT/speech-to-text) — file-based transcription reference
* [STT Realtime API](/vachana/STT/stt-websocket) — WebSocket protocol reference
* [TTS REST API](/vachana/TTS/tts-inference) — synchronous synthesis reference
* [TTS Realtime API](/vachana/TTS/tts-websocket) — WebSocket TTS reference
* [`gnani-vachana` on PyPI](https://pypi.org/project/gnani-vachana/) — core SDK
* [`pipecat-gnani` on PyPI](https://pypi.org/project/pipecat-gnani/) — this plugin
* [Pipecat Docs](https://docs.pipecat.ai/) — Pipecat framework reference
# New file
Source: https://docs.gnani.ai/temo
Description of your new file.
**Call Settings:**
* **Pre-Call Variables:** Define key details before the call, making conversations more personalized and efficient. Your agent starts the call with relevant details like customer name, account status, issue type, etc.
* Add as many variables as required and specify the details:
* **Unique Identifier:** Select one variable as the primary key for each call.
* **Variable:** A named placeholder storing a value from a database (e.g., *customer\_name, policy\_status*).
* **Description:** A short note for the LLM to understand the variable’s purpose.
* Using Pre-Call Variables:
* In the **Greeting Message, Ending Message,** and **System Prompt**, use double curly braces to insert variables dynamically.
* Example: *"Hi , I'm speaking from Gnani Technical Support."*
* *Example:* If the agent is calling a customer regarding their insurance policy, pre-call variables like *policy\_status* and *renewal\_date* ensure it starts with relevant information instead of asking redundant questions.
*And just like that, your GenAI agent is born! 🎉*
### **Lesson 5: Categorizing Call Outcomes**
Call Dispositions help you track and categorize the outcomes of your agent’s calls. Located within **Analytics Config** tab, this feature allows you to define custom statuses, enabling structured call analysis. To set up Call Dispositions:
* Toggle the **Call Dispositions** feature ON and enter a **default prompt** that helps the LLM understand how to categorize calls.
* Click **Add Disposition** to create a new category. Each disposition includes:
* [**Prompt:**](/M8_Disposition) Define when a call should be classified under this status.
* Once configured, the AI will automatically categorize calls based on the provided conditions and you can see the predefined disposition in Call Analytics and Agent Analytics.
### **Lesson 6: Testing your Agent**
**Testing via Chat:**
* Click on **Test → Chat Window → Start Testing \>** to use the built-in chat interface to test responses.
**Testing Voice Interactions:**
* Click on **Test → Web-based (Voice) → Start Testing \>** to use the web-based voice test to simulate real-world scenarios.
* To share the web-based voice testing tool, go to **Test → Web-based (Voice) → Generate Sharable Link**. The link works for 5 minutes, perfect for giving teammates or clients quick access without login requirements.
***
*Pro Tip:* Experiment with different temperatures, system prompts, transcribers and text-to-speech to see what best suits your use case. The right balance can make your agent more engaging, effective and tailored.
# Customise your Analytics
Source: https://docs.gnani.ai/untitled-page
The **Analytics Settings** section allows administrators to configure how call data is analysed and interpreted.
These configurations define what insights are extracted from conversations and how they appear across analytics dashboards and call-level analysis.
Administrators can configure:
* **Key words and phrase matching** – Detect specific words or phrases in conversations
* **Key Metrics** – Categorize calls based on predefined business conditions
* **Call Quality** – Monitor call quality signals such as silence
* **AI Call Analytics** – Extract deeper insights from conversations using AI
Once configured, these analytics are automatically applied to **all processed calls**.
Configured metrics and indicators are visible in:
* **Dashboard analytics**
* **Call listing view**
* **Individual call analytics page**
## 1. Key words and phrase matching
**Key Indicators** are used to detect the presence of specific **keywords or phrases** within call transcripts.
These indicators help teams track important signals in conversations, such as compliance statements, payment mentions, or escalation phrases.
The system scans transcripts and identifies **exact or similar phrase matches**.
***
## 1.1 Adding a Key Indicator
To create a new indicator:
1. Navigate to **Org Settings → Analytics Settings**
2. Select **Key Indicators**
3. Click **Add Key Indicators**
4. Enter the required information
5. Click **Add**
***
## 1.2 Key Indicator Fields
### Key Indicator
The name of the indicator you want to track.
Example:
* Payment Mention
* Escalation
* Verification Completed
***
### Keywords
Keywords define the words or phrases that the system should detect in transcripts.Multiple keywords can be added.
Example:
```text theme={null}
paid
payment done
already paid
transaction completed
```
***
### Persona
Defines whose speech should be evaluated for the keyword.\
Options include:
* **Agent**
* **Customer**
* **Both**
***
## Where Key Indicators Appear
Once configured, key indicators appear in:
* **Dashboard analytics**
* **Individual call analysis**
They help quickly identify whether important phrases occurred during the conversation
## 2. Key Metrics ( Call Classification)
**Key Metrics** are used to **categorize calls based on defined business conditions**.\
These metrics help classify the intent or outcome of a call.
For example:
* Payment Claimed
* Complaint Raised
* Information Request
Key metrics allow teams to analyze conversations at scale and understand the most common call outcomes.
***
### 2.1 Adding a Key Metric
To create a key metric:
1. Navigate to **Analytics Settings**
2. Select **Key Metrics**
3. Click **Add Key Metrics**
4. Enter the metric name and prompt
5. Click **Add**
***
### 2.2 Key Metric Fields
### Key Metric
The category name used to classify calls.
Examples:
* Payment Claimed
* Payment Pending
* Dispute Raised
***
### Prompt
The prompt defines the logic used to classify the call.\
It describes the scenario the system should detect.
Example:
```text theme={null}
Use this category when the customer clearly states that payment has already been completed.
```
The system analyzes the transcript and assigns the category if the condition is met.
***
## Where Key Metrics Appear
Configured key metrics appear in:
* **Call listing table (Key Metrics column)**
* **Dashboard analytics**
* **Single call analytics view**
They allow teams to quickly understand **why the call happened or what the outcome was**.
## 3. Call Quality
The **Call Quality** configuration allows administrators to define thresholds for identifying call quality issues.\
Currently, call quality analysis focuses on **silence detection**.\
Calls with excessive silence can be automatically flagged for review.
***
### 3.1 Configuring Silence Threshold
The **Dead Air Threshold** defines the maximum allowed silence duration within a call.\
If the silence exceeds this duration, the call may be flagged as a potential quality issue.
***
### 3.2 Steps to Configure
1. Navigate to **Analytics Settings**
2. Select **Call Quality**
3. Adjust the **Dead Air Threshold**
4. Save the configuration
***
### 3.3 Example
If the threshold is set to:
```text theme={null}
10 seconds
```
Any silence longer than 10 seconds will be considered **dead air**.
***
### 3.4 Where Call Quality Appears
Call quality insights appear in:
* **Call analytics page - Advanced call metrics**
## 4. AI Call Analytics
**AI Call Analytics** enables administrators to define **custom AI-powered metrics** that extract specific information from conversations.
These metrics allow organizations to automatically analyze calls for structured insights.
Examples include:
* Customer intent
* Call resolution status
* Issue type
* Customer sentiment patterns
***
### 4.1 Adding an AI Analytics Metric
To create a new AI analytics metric:
1. Navigate to **Analytics Settings**
2. Select **AI Call Analytics**
3. Click **Add Metrics**
4. Enter the metric name and prompt
5. Click **Add**
***
### 4.2 AI Analytics Fields
\
**Metric**
Name of the metric to extract.\
Examples:
* Call Outcome
* Customer Intent
* Escalation Risk
***
### Prompt
Defines what the AI should extract from the conversation.
Example:
```text theme={null}
Identify the main issue the customer is calling about.
```
The AI analyzes the conversation transcript and generates the required output.
***
### 4.3 Where AI Analytics Appear
Configured AI metrics appear in:
Call Analytics - Displayed in **individual call insights**.
# Manage Your Users
Source: https://docs.gnani.ai/untitled-page-copied-1
# Overview
The **Manage Users** section allows administrators to create, manage, and maintain users within the workspace. Users represent individuals who interact with the system, such as agents, supervisors, or administrators.
From this section, administrators can:
* Add new users
* Edit existing users
# Viewing Users
The **Manage Users** page displays a list of all users within the organization.
Each row in the table provides key information about a user.
### Information displayed
* **Name** – User’s full name
* **Email Address** – Associated email ID
* **Role** – Role assigned to the user (for example: Agent)
* **Status** – Indicates whether the user is active or inactive
Administrators can also perform actions such as editing user details directly from this page.
# Adding Users
New users can be added to the system individually or in bulk.
To add users:
1. Navigate to **Org Settings → Manage Users**
2. Click **Add User(s)**
3. Select the method for adding users
Available options include:
* **Single User**
* **Multiple Users**
***
## Adding a Single User
Single user creation allows administrators to manually add one user at a time.
### Steps
1. Click **Add User(s)**
2. Select **Single User**
3. Click **Continue**
4. Enter the required user details
5. Click **Create**
***
### User Details
The following fields are required when creating a user.
### Name
First name and last name of the user.
### Unique ID
A unique identifier assigned to the user within the system.
### Role
Defines the responsibilities and permissions assigned to the user.
Example roles may include:
* Agent
* Supervisor
* Administrator
### Status
Defines whether the user account is active.
### Email Address
The user's email address used for identification and communication.
***
### Additional Information
Optional information can also be provided.
### Country
Country associated with the user.
### Timezone
Timezone used for displaying time-related data.
### Bio
Short description or introduction for the user.
***
### Uploading Additional Documents
Administrators can upload supporting documents related to the user.
Examples may include:
* Identification documents
* Role verification documents
Supported file formats include:
* PDF
* DOC
* DOCX
***
## Adding Multiple Users (Bulk Upload)
Multiple users can be created simultaneously using a bulk upload feature.
This is useful when onboarding large teams.
***
### Steps to Add Multiple Users
1. Click **Add User(s)**
2. Select **Multiple Users**
3. Click **Continue**
4. Upload a user data file
Supported file formats include:
* CSV
* XLS
* XLSX
***
### Uploading the User File
Administrators can upload a spreadsheet containing user details.
The file should include required user information, such as:
* Name
* Email address
* Unique ID
* Role
A **sample file template** is available for download to ensure the correct file format.
Once uploaded, the system processes the file and creates users in bulk.
***
## Editing Users
Existing users can be updated at any time.
### Steps
1. Navigate to **Manage Users**
2. Locate the user in the list
3. Click the **Edit** option
4. Update the required information
5. Save the changes
***
### Updating User Information
Administrators can update:
* User name
* Email address
* Role
* Status
* Additional profile details
Changes take effect immediately after saving.
# Speech-to-Text (REST)
Source: https://docs.gnani.ai/vachana/STT/speech-to-text
POST /stt/v3
Quick transcription of audio clips up to 60 seconds via HTTP.
## Overview
The REST endpoint transcribes an audio file in a single synchronous HTTP request and returns the transcript immediately. It is best suited for short, pre-recorded audio clips.
| Use case | Recommended endpoint |
| --------------------------------- | ------------------------------------------------------ |
| Short clips ≤ 60 s (ideal ≤ 30 s) | **This endpoint** |
| Live microphone / real-time audio | [STT Realtime (WebSocket)](/vachana/STT/stt-websocket) |
| Large files or bulk jobs | [STT Batch](/vachana/STT/stt-batch) |
***
## Endpoint
```text theme={null}
POST https://api.vachana.ai/stt/v3
Content-Type: multipart/form-data
```
***
## Authentication
Pass your API key in the request header.
| Header | Type | Required | Description |
| -------------- | -------- | -------- | ------------------------------------------------------------------------- |
| `X-API-Key-ID` | `string` | ✅ | Your Gnani Prisma v2.5 API key. Obtain one from the Gnani APIs dashboard. |
***
## Request Parameters
All parameters are sent as `multipart/form-data` fields.
Audio file to transcribe. Supported formats: WAV, MP3, OGG, FLAC, AAC, M4A. Maximum duration: 60 seconds (ideal ≤ 30 s).
BCP-47 language code. See [Supported Languages](#supported-languages) below. Pass a comma-separated list of codes to enable auto-detection.
Forces processing with the single-language model for the specified code. Must be one of the values passed in `language_code`. Useful to improve accuracy when the audio is predominantly one language.
`verbatim` — raw spoken-form output. `transcribe` — enables Inverse Text Normalization (ITN): numbers, currency, dates, and phone numbers are written in their conventional form. See [ITN](#inverse-text-normalization-itn) below.
When `format=transcribe`, set `true` to render digits in the native script of the target language (e.g. `₹५,०००` instead of `₹5,000` for Hindi). Has no effect when `format=verbatim`. Currently supported for `hi-IN` and `en-IN` only.
***
## Response
### 200 — Success
```json theme={null}
{
"success": true,
"request_id": "req_abc123",
"timestamp": "20251226_143052.123",
"transcript": "नमस्ते, आप कैसे हैं?"
}
```
| Field | Type | Description |
| ------------ | --------- | --------------------------------------------------------------------------------------- |
| `success` | `boolean` | `true` when transcription completed without error. |
| `request_id` | `string` | Unique identifier for this request. Use it when contacting support or correlating logs. |
| `timestamp` | `string` | Server-side request timestamp in `YYYYMMDD_HHMMSS.mmm` format. |
| `transcript` | `string` | The transcribed text. Format depends on the `format` parameter. |
### Error Responses
| Status | Meaning |
| ------ | ------------------------------------------------------------------------ |
| `400` | Bad request — invalid parameters or unsupported audio format. |
| `429` | Rate limit exceeded — slow down or contact support to increase limits. |
| `500` | Internal server error — transient issue on our side; retry with backoff. |
| `503` | Service unavailable — the STT service is temporarily down. |
***
## Code Example
```bash cURL theme={null}
curl --request POST \
--url https://api.vachana.ai/stt/v3 \
--header 'Content-Type: multipart/form-data' \
--header 'X-API-Key-ID: ' \
--form audio_file='@recording.wav' \
--form language_code=hi-IN \
--form format=transcribe \
--form itn_native_numerals=true
```
```python Python SDK theme={null}
from gnani.stt import GnaniSTTClient
client = GnaniSTTClient(
organization_id="your-organization-id",
api_key="your-api-key",
user_id="your-user-id",
)
result = client.transcribe("recording.wav", language_code="hi-IN")
print(result["transcript"])
```
***
## Python SDK
The official Python SDK handles multipart construction, authentication headers, and retries automatically.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
The client requires three credentials: `organization_id`, `api_key`, and `user_id`. You can pass them directly or load them from environment variables.
```python Constructor arguments theme={null}
from gnani.stt import GnaniSTTClient
client = GnaniSTTClient(
organization_id="your-organization-id",
api_key="your-api-key",
user_id="your-user-id",
)
```
```bash Environment variables theme={null}
export GNANI_ORGANIZATION_ID="your-organization-id"
export GNANI_API_KEY="your-api-key"
export GNANI_USER_ID="your-user-id"
```
```python Environment variables (usage) theme={null}
from gnani.stt import GnaniSTTClient
# Picks up credentials from environment automatically
client = GnaniSTTClient()
```
### Transcribe Audio
```python From a file path theme={null}
result = client.transcribe("recording.wav", language_code="hi-IN")
print(result["transcript"])
```
```python From a file object theme={null}
with open("recording.wav", "rb") as f:
result = client.transcribe(f, language_code="hi-IN")
print(result["transcript"])
```
```python From raw bytes theme={null}
with open("recording.wav", "rb") as f:
audio_bytes = f.read()
result = client.transcribe(audio_bytes, language_code="hi-IN")
print(result["transcript"])
```
### Custom Request ID
Pass a `request_id` to correlate SDK calls with your own logs or support tickets.
```python theme={null}
result = client.transcribe(
"call.flac",
language_code="hi-IN",
request_id="my-trace-123",
)
```
### Error Handling
```python theme={null}
from gnani.stt import (
AuthenticationError,
InvalidAudioError,
APIError,
)
try:
result = client.transcribe("audio.wav", language_code="hi-IN")
print(result["transcript"])
except AuthenticationError:
print("Invalid credentials — check your organization_id, api_key, and user_id.")
except InvalidAudioError as e:
print(f"Bad audio file: {e}")
except APIError as e:
print(f"API error {e.status_code}: {e}")
```
***
## Supported Languages
The Gnani Prisma v2.5 API supports 10 Indian languages.
| Language | Code | Native Script | Example |
| --------- | ------- | ------------------- | ------------------------------- |
| Bengali | `bn-IN` | Bengali (বাংলা) | "আমি ভাত খাই" |
| English | `en-IN` | Latin | "I am going to the market" |
| Gujarati | `gu-IN` | Gujarati (ગુજરાતી) | "હું બજાર જાઉં છું" |
| Hindi | `hi-IN` | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
| Kannada | `kn-IN` | Kannada (ಕನ್ನಡ) | "ನಾನು ಮಾರುಕಟ್ಟೆಗೆ ಹೋಗುತ್ತೇನೆ" |
| Malayalam | `ml-IN` | Malayalam (മലയാളം) | "ഞാൻ ചന്തയിലേക്ക് പോകുന്നു" |
| Marathi | `mr-IN` | Devanagari (मराठी) | "मी बाजारात जातोय" |
| Punjabi | `pa-IN` | Gurmukhi (ਪੰਜਾਬੀ) | "ਮੈਂ ਬਾਜ਼ਾਰ ਜਾ ਰਿਹਾ ਹਾਂ" |
| Tamil | `ta-IN` | Tamil (தமிழ்) | "நான் சந்தைக்கு செல்கிறேன்" |
| Telugu | `te-IN` | Telugu (తెలుగు) | "నేను మార్కెట్కి వెళ్తున్నాను" |
For **auto-detection**, pass the full set of candidate language codes as a comma-separated value in `language_code`. For example: `en-IN,hi-IN,ta-IN,te-IN,kn-IN,ml-IN,gu-IN,mr-IN,bn-IN,pa-IN`.
***
## Inverse Text Normalization (ITN)
ITN converts the spoken-form output of the ASR engine into the conventional written form a reader expects — numbers become digits, currency gets the ₹ symbol, dates are formatted, and phone numbers are compacted — all in one pass, immediately after transcription.
**How to enable:** Set `format=transcribe` in the request body.
Currently supported for **Hindi (`hi-IN`)** and **English (`en-IN`)** only. All other languages use `verbatim` output regardless of the `format` value.
### What ITN Normalizes
#### 1 — Cardinal & Ordinal Numbers
Whole numbers and positional ranks are formatted using Indian comma grouping (groups of 2 after the first 3 digits).
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------- | -------------------- | ----------------------- |
| दो हज़ार | 2,000 | Indian comma grouping |
| पाँच लाख बीस हज़ार | 5,20,000 | Lakh-scale grouping |
| उन्नीस सौ चौरानवे | 1,994 | Hundred-base year form |
| five lakh | 5,00,000 | English lakh convention |
| पहला / twenty first | 1st / 21st | Ordinal suffix |
#### 2 — Currency & Money
All Indian currency expressions — including paise fractions and lakh/crore scales — are formatted with the ₹ symbol and Indian comma grouping.
| Spoken input (ASR) | Written output (ITN) | Rule |
| --------------------------- | -------------------- | ---------------------- |
| पाँच सौ रुपये | ₹500 | ₹ + amount |
| तीन रुपये पचास पैसे | ₹3.50 | ₹ + rupees.paise |
| दस लाख रुपये | ₹10,00,000 | ₹ + lakh grouping |
| I need five thousand rupees | ₹5,000 | English India pipeline |
#### 3 — Dates
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------ | -------------------- | ----------------------- |
| बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 | DD Month YYYY (hi) |
| fifteenth january twenty twenty five | 15th January 2025 | Ordinal Month YYYY (en) |
#### 4 — Times
Indian time-of-day words (सुबह, दोपहर, शाम, रात) automatically map to 24-hour HH:MM output.
| Spoken input (ASR) | Written output (ITN) | Rule |
| -------------------------------------- | ---------------------------- | ----------------------- |
| सुबह पाँच बजे | सुबह 05:00 | सुबह = AM |
| शाम पाँच बजे | शाम 17:00 | शाम = evening (16–20 h) |
| रात के दस बजे | रात 22:00 | रात = night (20–24 h) |
| meeting at five fifteen in the evening | meeting 17:15 in the evening | en — 24-hour |
#### 5 — Phone Numbers & PIN Codes
Digit streams are concatenated into compact numeric strings. 10-digit streams → mobile number; 6-digit streams → PIN. Repeat prefixes (double/डबल, triple/ट्रिपल) are expanded.
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------- | -------------------- | ------------------- |
| नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 | 10 digits → phone |
| एक एक शून्य शून्य शून्य एक | 110001 | 6 digits → PIN |
| डबल आठ नौ शून्य एक दो तीन चार पाँच छह | 8890123456 | double prefix |
| one two three four five six | 123456 | English digit words |
#### 6 — Mixed & Code-Mixed Utterances
A single sentence may contain multiple entity types or blend Hindi and English. ITN handles all in one pass, normalizing each entity independently.
| Spoken input (ASR) | Written output (ITN) |
| ----------------------------------------------------- | --------------------------------- |
| कल थ्री फिफ्टी पीएम को पाँच सौ रुपये transfer करना है | कल 15:50 को ₹500 transfer करना है |
### Native Script Digits — `itn_native_numerals`
By default, ITN outputs Western Arabic digits (0–9) regardless of language. Set `itn_native_numerals=true` to render digits in the native script of the target language.
| Language | Spoken input | `false` (default) | `true` — native script |
| --------------- | -------------------- | ----------------- | -------------------------- |
| Hindi `hi-IN` | पाँच हज़ार रुपये | ₹5,000 | ₹५,००० |
| English `en-IN` | five thousand rupees | ₹5,000 | ₹5,000 (Latin — no change) |
### What ITN Does Not Change
ITN intentionally preserves idiomatic and ambiguous phrases to avoid incorrect normalization.
* **दो तीन** (meaning *a few*) stays as text, not `2` or `3`
* **कर दो / ले दो** (imperative verbs) are kept as words, not treated as cardinal 2
If a word or phrase is unchanged in the output, treat it as a failure only when the input was unambiguously a numeric entity.
# Speech-to-Text (Batch)
Source: https://docs.gnani.ai/vachana/STT/stt-batch
POST /stt/v3/batch/submit
Asynchronous transcription of long or multiple audio files via HTTP.
## Overview
Submit one or more audio files for transcription and receive a `job_id` immediately. Poll the status endpoint on a fixed interval until the job completes and transcripts are available. Ideal for long recordings, bulk files, or offline pipelines where you do not need a live response. For real-time transcription, see [STT Realtime](/vachana/STT/stt-websocket). For short clips under 60 seconds, see [STT REST](/vachana/STT/speech-to-text).
## Endpoints
| Operation | Method | URL |
| ---------------- | ------ | ----------------------------------------------------- |
| Submit Job | POST | `https://api.vachana.ai/stt/v3/batch/submit` |
| Check Job Status | GET | `https://api.vachana.ai/stt/v3/batch/status/{job_id}` |
## Limits & Specifications
| Item | Limit |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| Max audio duration | Less than **1 hour** per file |
| Max files per request | **10 files** per API call |
| Max total payload size | **80 MB** across all files and form fields combined |
| Minimum poll interval | **60 seconds** between status calls for the same `job_id` |
| Speaker diarization. | This API supports **at most 2 speakers** per file (two-party diarization). Scenarios with more than two distinct speakers are not supported |
### Supported Audio Formats
`AAC` · `WAV` · `FLAC` · `ALAC` · `OGG (Vorbis)` · `Opus`
Use standard file extensions and MIME types (e.g. `.m4a` for AAC, `.wav`, `.flac`, `.ogg`).
## Authentication
Send these headers on **every request** both submit and status calls.
| Header | Required | Description |
| ------------------ | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| `X-API-Key-ID` | Yes | Your API key. Required for all requests. |
| `X-API-Request-ID` | No | A unique trace ID (e.g. UUID) you assign. Used to correlate your logs with platform logs or support. If omitted, the platform may generate one. |
Do **not** set `Content-Type: application/json` on the submit request. Use `multipart/form-data`. curl sets the correct boundary automatically when you use `-F` / `--form`.
***
## Submit a Transcription Job
### `POST /stt/v3/batch/submit`
Upload audio files and kick off an asynchronous transcription job. The response returns a `job_id` immediately. the files are not yet transcribed at this point.
#### Request — Form Fields
#### Supported Language Codes
| Language | Code | Native Script | Example Text |
| --------- | ------- | ------------------- | ------------------------------- |
| Bengali | `bn-IN` | Bengali (বাংলা) | "আমি ভাত খাই" |
| English | `en-IN` | Latin | "I am going to the market" |
| Gujarati | `gu-IN` | Gujarati (ગુજરાતી) | "હું બજાર જાઉં છું" |
| Hindi | `hi-IN` | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
| Kannada | `kn-IN` | Kannada (ಕನ್ನಡ) | "ನಾನು ಮಾರುಕಟ್ಟೆಗೆ ಹೋಗುತ್ತೇನೆ" |
| Malayalam | `ml-IN` | Malayalam (മലയാളം) | "ഞാൻ ചന്തയിലേക്ക് പോകുന്നു" |
| Marathi | `mr-IN` | Devanagari (मराठी) | "मी बाजारात जातोय" |
| Punjabi | `pa-IN` | Gurmukhi (ਪੰਜਾਬੀ) | "ਮੈਂ ਬਾਜ਼ਾਰ ਜਾ ਰਿਹਾ ਹਾਂ" |
| Tamil | `ta-IN` | Tamil (தமிழ்) | "நான் சந்தைக்கு செல்கிறேன்" |
| Telugu | `te-IN` | Telugu (తెలుగు) | "నేను మార్కెట్కి వెళ్తున్నాను" |
#### Example — curl
```bash theme={null}
curl --location --request POST 'https://api.vachana.ai/stt/v3/batch/submit' \
--header 'X-API-Key-ID: ' \
--header 'X-API-Request-ID: 550e8400-e29b-41d4-a716-446655440000' \
--form 'language_code=hi-IN' \
--form 'is_multi_channel=false' \
--form 'format=transcribe' \
--form 'audio_files=@"/path/to/first.wav"' \
--form 'audio_files=@"/path/to/second.wav"'
```
#### Response — `200 OK`
```json theme={null}
{
"job_id": "batch_7f3a92c1d4e8",
"status": "submitted",
"file_count": 2,
"message": "Job accepted. Poll the status endpoint every 60 seconds for results."
}
```
| Field | Type | Description |
| ------------ | ------- | -------------------------------------------------- |
| `job_id` | string | Identifier for this job. Use it in the status URL. |
| `status` | string | Initial value is always `submitted`. |
| `file_count` | integer | Number of files accepted into the job. |
| `message` | string | Short confirmation with polling instructions. |
#### Errors
| HTTP Status | When |
| ----------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `400` | No files uploaded, empty file, more than 10 files, payload over 80 MB, unsupported format, or other client-side validation failure. |
| `500` | Server error. |
***
## Check Job Status
### `GET /stt/v3/batch/status/{job_id}`
Poll this endpoint to check progress and retrieve transcription results once the job finishes. Call this **once every 60 seconds** per `job_id`. do not poll more frequently.
#### Path Parameter
| Parameter | Required | Description |
| --------- | -------- | ----------------------------------------------- |
| `job_id` | Yes | The `job_id` returned from the Submit response. |
#### Example — curl
```bash theme={null}
curl --location --request GET 'https://api.vachana.ai/stt/v3/batch/status/{job_id}' \
--header 'X-API-Key-ID: ' \
--header 'X-API-Request-ID: 550e8400-e29b-41d4-a716-446655440000'
```
#### Response — `200 OK`
```json theme={null}
{
"job_id": "batch_7f3a92c1d4e8",
"status": "completed",
"total_files": 2,
"completed_files": 2,
"failed_files": 0,
"overall_progress": 100,
"error": null,
"results": [
{
"filename": "first.wav",
"status": "completed",
"full_transcript": "नमस्ते, आप कैसे हैं?",
"total_duration": 45.3,
"error": null,
"segments": [
{
"segment_id": 0,
"start_time": 0.0,
"end_time": 3.2,
"text": "नमस्ते, आप कैसे हैं?",
"speaker_id": 1,
"language_detected": "hi-IN"
}
]
}
]
}
```
#### Job-Level Response Fields
| Field | Type | Description |
| ------------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| `job_id` | string | Job identifier. |
| `status` | string | `submitted` — accepted or in progress. `processing` — actively transcribing. `completed` — done. `failed` — job-level failure. |
| `total_files` | integer | Total number of files in the job. |
| `completed_files` | integer | Files finished successfully. Meaningful only when the job has reached a final state. |
| `failed_files` | integer | Files that failed. Meaningful only when the job has reached a final state. |
| `overall_progress` | integer | Approximate progress from `0` to `100` while the job is running. |
| `results` | array or `null` | Per-file results. `null` while the job is `submitted` or `processing`. |
| `error` | string or `null` | Top-level error message for the job, if any. |
#### Per-File Result Fields — `results[]`
| Field | Type | Description |
| ----------------- | ---------------- | ------------------------------------------------- |
| `filename` | string | Original file name as submitted. |
| `full_transcript` | string | Complete transcribed text for the file. |
| `segments` | array | Time-aligned transcript segments (see below). |
| `total_duration` | number | Audio duration in seconds. |
| `status` | string | `completed` or `failed` for this individual file. |
| `error` | string or `null` | Error message for this file if it failed. |
#### Per-Segment Fields — `results[].segments[]`
| Field | Type | Description |
| ------------------- | ------- | ------------------------------------------------------ |
| `segment_id` | integer | Segment index (zero-based). |
| `start_time` | number | Segment start time in seconds. |
| `end_time` | number | Segment end time in seconds. |
| `text` | string | Transcribed text for this segment. |
| `speaker_id` | integer | Speaker identifier. Populated for multi-channel audio. |
| `language_detected` | string | BCP-47 code of the detected language for this segment. |
#### Errors
| HTTP Status | When |
| ----------- | ------------------------------------------------------------------ |
| `404` | `job_id` not found — unknown ID or the job is no longer available. |
| `500` | Server error. |
***
## Inverse Text Normalization (ITN)
When `format=transcribe` is passed in the form body, ITN runs on every file's transcript after recognition — converting spoken-form numbers, currency, dates, times, and phone numbers into the compact written form a reader expects.
ITN is currently supported for **Hindi (`hi-IN`)** and **English (`en-IN`)** only. Enabling ITN for other languages has no effect — transcripts are returned as-is.
#### What ITN Normalizes
ITN recognizes six categories of spoken expressions. Every matching span is transformed; all other words pass through unchanged.
Whole numbers and positional ranks are formatted using Indian comma grouping (groups of 2 after the first 3 digits).
| Spoken input (ASR) | Written output (ITN) | Format rule |
| ------------------- | -------------------- | ----------------------- |
| दो हज़ार | 2,000 | Indian comma grouping |
| पाँच लाख बीस हज़ार | 5,20,000 | Lakh-scale grouping |
| five lakh | 5,00,000 | English lakh convention |
| पहला / twenty first | 1st / 21st | Ordinal suffix |
All Indian currency expressions — including paise fractions and lakh/crore scales — are formatted with the ₹ symbol and Indian comma grouping.
| Spoken input (ASR) | Written output (ITN) | Format rule |
| --------------------------- | -------------------- | ---------------------- |
| पाँच सौ रुपये | ₹500 | ₹ + amount |
| तीन रुपये पचास पैसे | ₹3.50 | ₹ + rupees.paise |
| I need five thousand rupees | ₹5,000 | English India pipeline |
| pay do lakh rupees | ₹2,00,000 | Code-mixed en/hi |
| Spoken input (ASR) | Written output (ITN) | Format rule |
| ------------------------------------ | -------------------- | ----------------------- |
| बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 | DD Month YYYY (hi) |
| fifteenth january twenty twenty five | 15th January 2025 | Ordinal Month YYYY (en) |
Indian time-of-day words (सुबह, दोपहर, शाम, रात) automatically map to 24-hour HH:MM output.
| Spoken input (ASR) | Written output (ITN) | Format rule |
| -------------------------------------- | ---------------------------- | ---------------------- |
| सुबह पाँच बजे | सुबह 05:00 | सुबह = AM |
| शाम पाँच बजे | शाम 17:00 | शाम = evening (16–20h) |
| रात के दस बजे | रात 22:00 | रात = night (20–24h) |
| meeting at five fifteen in the evening | meeting 17:15 in the evening | en — 24-hour |
10-digit streams → mobile number; 6-digit streams → PIN. Repeat prefixes (double/डबल) are expanded.
| Spoken input (ASR) | Written output (ITN) | Format rule |
| ------------------------------------- | -------------------- | ------------------- |
| नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 | 10 digits → phone |
| एक एक शून्य शून्य शून्य एक | 110001 | 6 digits → PIN |
| one two three four five six | 123456 | English digit words |
A single file may contain segments with multiple entity types or blend Hindi and English. ITN normalizes each entity independently in one pass.
| Spoken input (ASR) | Written output (ITN) |
| ----------------------------------------------------- | --------------------------------- |
| कल थ्री फिफ्टी पीएम को पाँच सौ रुपये transfer करना है | कल 15:50 को ₹500 transfer करना है |
| pay do lakh rupees by fifteenth march | pay ₹2,00,000 by 15th March |
#### Native Script Digits — `itn_native_numerals`
By default, ITN outputs Western Arabic digits (0–9) regardless of language. When `format=transcribe` is set, you can additionally pass `itn_native_numerals=true` to render digits in the native script of the target language.
| Language | Spoken input | `false` (default) | `true` — native script |
| --------------- | -------------------- | ----------------- | -------------------------- |
| Hindi `hi-IN` | पाँच हज़ार रुपये | ₹5,000 | ₹५,००० |
| English `en-IN` | five thousand rupees | ₹5,000 | ₹5,000 (Latin — no change) |
English always outputs Western Arabic digits. `itn_native_numerals=true` has no effect for `en-IN`.
#### What ITN Does Not Change
ITN intentionally preserves idiomatic and ambiguous phrases.
* **दो तीन** (meaning *a few*) stays as text, not `2` or `3`
* **कर दो / ले दो** (imperative verbs) are kept as words, not treated as cardinal 2
***
## Flow Summary
1. **Submit** — `POST https://api.vachana.ai/stt/v3/batch/submit` with `X-API-Key-ID`, optional `X-API-Request-ID`, and form fields `audio_files`, `language_code`, and optionally `is_multi_channel`, `format`, and `itn_native_numerals`.
2. **Save the `job_id`** from the submit response.
3. **Poll** — `GET https://api.vachana.ai/stt/v3/batch/status/{job_id}` (same auth headers) **every 60 seconds** until `status` is `completed` or `failed` and `results` is populated.
# Speech-to-Text (Realtime)
Source: https://docs.gnani.ai/vachana/STT/stt-websocket
Real-time speech-to-text over a persistent WebSocket connection.
## Overview
Stream raw PCM audio frames and receive transcript segments as speech is detected. The server uses Voice Activity Detection (VAD) to identify speech boundaries and returns a transcript for each segment.
| Use case | Recommended endpoint |
| ---------------------------------------------- | --------------------------------------- |
| Live microphone / phone call / real-time audio | **This endpoint** |
| Short pre-recorded clips ≤ 60 s | [STT REST](/vachana/STT/speech-to-text) |
| Large files or bulk jobs | [STT Batch](/vachana/STT/stt-batch) |
***
## Endpoint
```text theme={null}
WSS wss://api.vachana.ai/stt/v3/stream
```
***
## Connection Headers
All configuration is passed as WebSocket upgrade headers at connection time. Headers cannot be changed mid-session — reconnect with new headers to change settings.
| Header | Required | Default | Description |
| --------------------- | -------- | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| `x-api-key-id` | ✅ | — | Your Gnani API key. |
| `lang_code` | ✅ | `en-IN` | BCP-47 language code for transcription. See [Supported Languages](#supported-languages). |
| `x-sample-rate` | ❌ | `16000` | Sample rate of the audio stream in Hz. Accepted values: `8000`, `16000`, `44100`, `48000`. Must match the actual sample rate of your audio source. |
| `x-format` | ❌ | `verbatim` | `verbatim` — raw spoken-form output. `transcribe` — enables Inverse Text Normalization (ITN). See [ITN](#inverse-text-normalization-itn). |
| `itn_native_numerals` | ❌ | `false` | When `x-format=transcribe`, set `true` to render digits in the native script of the target language (e.g. `₹५,०००` instead of `₹5,000` for Hindi). |
**Choosing the right sample rate:**
| Value | When to use |
| ------- | ------------------------------------------------- |
| `48000` | Browser `getUserMedia` default; Mac microphone |
| `44100` | Mac microphone alternate; CD-quality audio |
| `16000` | Wideband telephony; sent as-is with no resampling |
| `8000` | Narrowband telephony (legacy PSTN / VoIP) |
***
## Connection Flow
A WebSocket session follows a strict sequence:
1. **Client connects** — opens a WebSocket to `/stt/v3/stream` with all required headers.
2. **Server confirms** — immediately sends a `connected` message echoing the active configuration.
3. **Client streams audio** — continuously sends binary frames of raw PCM audio at a steady real-time cadence.
4. **Server detects speech** — VAD identifies end-of-speech boundaries and emits a `processing` message to acknowledge that a segment was captured.
5. **Server returns transcript** — sends a `transcript` message with the transcribed text, segment metadata, and latency.
6. **Either side closes** — client or server may close the connection at any time.
The `processing` message is a low-latency signal that audio was captured and transcription has begun. Expect a `transcript` message shortly after.
***
## Audio Format & Sending Audio
All audio must be sent as **raw PCM binary frames** over the WebSocket. No container format (WAV, MP3, etc.) is accepted mid-stream.
### PCM Specification
| Property | 16 kHz | 8 kHz |
| ------------------- | --------------------------------------- | --------------------------------------- |
| Encoding | PCM signed 16-bit little-endian | PCM signed 16-bit little-endian |
| Sample Rate | 16,000 Hz | 8,000 Hz |
| Channels | 1 (mono) | 1 (mono) |
| Samples per chunk | 512 | 512 |
| **Bytes per frame** | **1,024 bytes** (512 samples × 2 bytes) | **1,024 bytes** (512 samples × 2 bytes) |
| Frame duration | 32 ms | 64 ms |
### Sending Rules
* Each binary frame must be **exactly 1,024 bytes**.
* Frames must be sent at **real-time cadence** — one frame every 32 ms (16 kHz) or 64 ms (8 kHz). Do not buffer and burst; this degrades VAD accuracy.
* For `44100` and `48000` Hz sources, the server resamples internally — still send 1,024-byte frames at the appropriate cadence.
***
## Server Messages
The server sends JSON text frames. All messages share a `type` discriminator field and an ISO-8601 `timestamp`.
### `connected`
Sent once immediately after the WebSocket handshake succeeds.
```json theme={null}
{
"type": "connected",
"message": "STT service ready — VAD service connected",
"timestamp": "2024-01-15T10:30:00.000Z",
"config": {
"sample_rate": 16000,
"chunk_size": 512
}
}
```
| Field | Type | Description |
| -------------------- | --------- | ------------------------------------------------------ |
| `type` | `string` | Always `"connected"`. |
| `message` | `string` | Human-readable status string. |
| `timestamp` | `string` | ISO-8601 server timestamp. |
| `config.sample_rate` | `integer` | Active sample rate in Hz, echoed from `x-sample-rate`. |
| `config.chunk_size` | `integer` | Expected chunk size in samples (always 512). |
### `processing`
Emitted when VAD detects the end of a speech segment and transcription has begun. Use this as a low-latency acknowledgment that audio was captured.
```json theme={null}
{
"type": "processing",
"timestamp": "2024-01-15T10:30:05.123Z"
}
```
| Field | Type | Description |
| ----------- | -------- | ------------------------------------------------ |
| `type` | `string` | Always `"processing"`. |
| `timestamp` | `string` | ISO-8601 timestamp when speech-end was detected. |
### `transcript`
Contains the transcribed text for a completed speech segment.
```json theme={null}
{
"type": "transcript",
"timestamp": "2024-01-15T10:30:05.987Z",
"text": "Hello, how are you today?",
"audio_duration_ms": 2340,
"segment_id": "",
"segment_index": 0,
"latency": 320
}
```
| Field | Type | Description |
| ------------------- | --------- | ---------------------------------------------------------------------------------------- |
| `type` | `string` | Always `"transcript"`. |
| `timestamp` | `string` | ISO-8601 timestamp when the transcript was emitted. |
| `text` | `string` | Transcribed text. Format depends on the `x-format` header. |
| `audio_duration_ms` | `integer` | Duration of the captured speech segment in milliseconds. |
| `segment_id` | `string` | Unique identifier for this speech segment. Use for deduplication or support correlation. |
| `segment_index` | `integer` | Sequential index of this segment within the session, starting at `0`. |
| `latency` | `integer` | Time in milliseconds from end-of-speech detection to transcript delivery. |
### `error`
Sent when the server encounters a recoverable or fatal error. The connection may remain open after a recoverable error.
| Field | Type | Description |
| ----------- | -------- | ---------------------------------------- |
| `type` | `string` | Always `"error"`. |
| `timestamp` | `string` | ISO-8601 timestamp of the error. |
| `message` | `string` | Human-readable description of the error. |
```json theme={null}
{
"type": "error",
"timestamp": "2024-01-15T10:30:10.000Z",
"message": "STT engine failed to initialize"
}
```
***
## Python SDK
The official Python SDK wraps the WebSocket connection, audio pacing, and event parsing into a clean async interface.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
The streaming client requires your API key and language code.
```python Constructor argument theme={null}
from gnani.stt import GnaniSTTStreamClient
stream = GnaniSTTStreamClient(
api_key="your-api-key",
language_code="hi-IN",
)
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.stt import GnaniSTTStreamClient
# Picks up GNANI_API_KEY from environment automatically
stream = GnaniSTTStreamClient(language_code="hi-IN")
```
### Stream Audio from a File
Use the async context manager and the `stream_audio` helper. It handles real-time pacing automatically so frames are sent at the correct cadence for VAD.
```python theme={null}
import asyncio
from gnani.stt import GnaniSTTStreamClient
async def main():
async with GnaniSTTStreamClient(
api_key="your-api-key",
language_code="hi-IN",
sample_rate=16000,
) as stream:
with open("audio.pcm", "rb") as f:
transcripts = await stream.stream_audio(
f,
on_transcript=lambda t: print(f"Transcript: {t.text}"),
on_processing=lambda p: print("Processing..."),
realtime_pace=True, # sends frames at real-time cadence
)
print(f"Total segments: {len(transcripts)}")
asyncio.run(main())
```
### Iterate Over Events Manually
For lower-level control — handling each event type differently or interleaving sending and receiving — iterate over the stream directly.
```python theme={null}
import asyncio
from gnani.stt import GnaniSTTStreamClient, StreamTranscriptEvent, StreamProcessingEvent
async def main():
async with GnaniSTTStreamClient(
api_key="your-api-key",
language_code="hi-IN",
) as stream:
with open("audio.pcm", "rb") as f:
while chunk := f.read(1024):
await stream.send_audio(chunk)
await asyncio.sleep(0.032) # 32 ms per frame at 16 kHz
async for event in stream:
if isinstance(event, StreamTranscriptEvent):
print(f"[Segment {event.segment_index}] {event.text}")
print(f" Duration: {event.audio_duration_ms} ms Latency: {event.latency} ms")
elif isinstance(event, StreamProcessingEvent):
print("Processing speech...")
asyncio.run(main())
```
### Using 8 kHz Audio (Telephony)
```python theme={null}
stream = GnaniSTTStreamClient(
api_key="your-api-key",
language_code="en-IN",
sample_rate=8000,
)
```
### SDK Event Types
All events are typed dataclasses with a `raw` field containing the full server JSON.
| Event class | Key fields | Description |
| ----------------------- | ------------------------------------------------------- | -------------------------------------------------------------------- |
| `StreamConnectedEvent` | `message`, `sample_rate`, `chunk_size` | Sent once after the WebSocket handshake. Confirms the active config. |
| `StreamProcessingEvent` | `timestamp` | VAD detected end-of-speech; transcription has started. |
| `StreamTranscriptEvent` | `text`, `segment_index`, `audio_duration_ms`, `latency` | Completed transcript for a speech segment. |
| `StreamErrorEvent` | `message`, `timestamp` | Server-side error, recoverable or fatal. |
### Error Handling
```python theme={null}
from gnani.stt import (
StreamConnectionError, # Could not establish the WebSocket connection
StreamClosedError, # Attempted to send on an already-closed stream
StreamError, # Server returned an error message mid-session
)
try:
async with GnaniSTTStreamClient(api_key="your-api-key") as stream:
await stream.send_audio(chunk)
except StreamConnectionError as e:
print(f"Could not connect: {e}")
except StreamClosedError as e:
print(f"Stream was already closed: {e}")
except StreamError as e:
print(f"Server error: {e.message} (at {e.timestamp})")
```
***
## Supported Languages
| Language | Code | Native Script | Example |
| --------- | ------- | ------------------- | ------------------------------- |
| Bengali | `bn-IN` | Bengali (বাংলা) | "আমি ভাত খাই" |
| English | `en-IN` | Latin | "I am going to the market" |
| Gujarati | `gu-IN` | Gujarati (ગુજરાતી) | "હું બજાર જાઉં છું" |
| Hindi | `hi-IN` | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
| Kannada | `kn-IN` | Kannada (ಕನ್ನಡ) | "ನಾನು ಮಾರುಕಟ್ಟೆಗೆ ಹೋಗುತ್ತೇನೆ" |
| Malayalam | `ml-IN` | Malayalam (മലയാളം) | "ഞാൻ ചന്തയിലേക്ക് പോകുന്നു" |
| Marathi | `mr-IN` | Devanagari (मराठी) | "मी बाजारात जातोय" |
| Punjabi | `pa-IN` | Gurmukhi (ਪੰਜਾਬੀ) | "ਮੈਂ ਬਾਜ਼ਾਰ ਜਾ ਰਿਹਾ ਹਾਂ" |
| Tamil | `ta-IN` | Tamil (தமிழ்) | "நான் சந்தைக்கு செல்கிறேன்" |
| Telugu | `te-IN` | Telugu (తెలుగు) | "నేను మార్కెట్కి వెళ్తున్నాను" |
For **auto-detection**, pass all desired language codes comma-separated in the `lang_code` header. For example: `en-IN,hi-IN,ta-IN,te-IN,kn-IN,ml-IN,gu-IN,mr-IN,bn-IN,pa-IN`.
***
## Inverse Text Normalization (ITN)
When `x-format: transcribe` is set, ITN runs on every transcript segment immediately after recognition — converting spoken-form numbers, currency, dates, times, and phone numbers into the compact written form a reader expects.
Currently supported for **Hindi (`hi-IN`)** and **English (`en-IN`)** only. Enabling ITN for other languages has no effect; transcripts are returned verbatim.
### What ITN Normalizes
#### 1 — Cardinal & Ordinal Numbers
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------- | -------------------- | ----------------------- |
| दो हज़ार | 2,000 | Indian comma grouping |
| पाँच लाख बीस हज़ार | 5,20,000 | Lakh-scale grouping |
| five lakh | 5,00,000 | English lakh convention |
| पहला / twenty first | 1st / 21st | Ordinal suffix |
#### 2 — Currency & Money
| Spoken input (ASR) | Written output (ITN) | Rule |
| --------------------------- | -------------------- | ---------------------- |
| पाँच सौ रुपये | ₹500 | ₹ + amount |
| तीन रुपये पचास पैसे | ₹3.50 | ₹ + rupees.paise |
| I need five thousand rupees | ₹5,000 | English India pipeline |
| pay do lakh rupees | ₹2,00,000 | Code-mixed en/hi |
#### 3 — Dates
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------ | -------------------- | ----------------------- |
| बीस जनवरी दो हज़ार पच्चीस | 20 जनवरी 2025 | DD Month YYYY (hi) |
| fifteenth january twenty twenty five | 15th January 2025 | Ordinal Month YYYY (en) |
#### 4 — Times
Indian time-of-day words (सुबह, दोपहर, शाम, रात) automatically map to 24-hour HH:MM output.
| Spoken input (ASR) | Written output (ITN) | Rule |
| -------------------------------------- | ---------------------------- | ----------------------- |
| सुबह पाँच बजे | सुबह 05:00 | सुबह = AM |
| शाम पाँच बजे | शाम 17:00 | शाम = evening (16–20 h) |
| रात के दस बजे | रात 22:00 | रात = night (20–24 h) |
| meeting at five fifteen in the evening | meeting 17:15 in the evening | en — 24-hour |
#### 5 — Phone Numbers & PIN Codes
| Spoken input (ASR) | Written output (ITN) | Rule |
| ------------------------------------- | -------------------- | ------------------- |
| नौ आठ सात छह पाँच चार तीन दो एक शून्य | 9876543210 | 10 digits → phone |
| एक एक शून्य शून्य शून्य एक | 110001 | 6 digits → PIN |
| one two three four five six | 123456 | English digit words |
#### 6 — Mixed & Code-Mixed Utterances
| Spoken input (ASR) | Written output (ITN) |
| ----------------------------------------------------- | --------------------------------- |
| कल थ्री फिफ्टी पीएम को पाँच सौ रुपये transfer करना है | कल 15:50 को ₹500 transfer करना है |
| pay do lakh rupees by fifteenth march | pay ₹2,00,000 by 15th March |
### Native Script Digits — `itn_native_numerals`
By default, ITN outputs Western Arabic digits (0–9). Set `itn_native_numerals: true` in the connection headers to render digits in the native script of the target language.
| Language | Spoken input | `false` (default) | `true` — native script |
| --------------- | -------------------- | ----------------- | -------------------------- |
| Hindi `hi-IN` | पाँच हज़ार रुपये | ₹5,000 | ₹५,००० |
| English `en-IN` | five thousand rupees | ₹5,000 | ₹5,000 (Latin — no change) |
### What ITN Does Not Change
ITN intentionally preserves idiomatic and ambiguous phrases to avoid incorrect normalization.
* **दो तीन** (meaning *a few*) stays as text, not `2` or `3`
* **कर दो / ले दो** (imperative verbs) are kept as words, not treated as cardinal 2
***
# Text-to-Speech (REST)
Source: https://docs.gnani.ai/vachana/TTS/tts-inference
POST /api/v1/tts/inference
Synchronous text-to-speech with full audio returned in one response.
**Currently in beta.** You're on the priority waitlist and among the first to get access.
## Overview
Get the complete synthesized audio in one response. Best for downloads or batch processing. For streaming playback, see [TTS Streaming](/vachana/TTS/tts-sse) or [TTS Realtime](/vachana/TTS/tts-websocket).
Passing numbers, IDs, dates, or currency as raw strings causes mispronunciations. See the [Input Formatting Guide](/vachana/TTS/tts-input-formating) for correct formatting of phone numbers, account numbers, PINs, Aadhaar, vehicle registration numbers, GSTIN, currency, and more.
***
## Available Voices
| Voice | Gender | Description |
| ------- | ------ | ------------------------ |
| Pranav | Male | Bold, Trustworthy |
| Kaveri | Female | Confident, Bright |
| Shubhra | Female | Gentle, Expressive |
| Deepak | Male | Grounded, Conversational |
***
## Python SDK
The official Python SDK lets you synthesize speech in one line, without constructing JSON payloads or handling binary audio responses manually.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
The TTS client requires only your API key.
```python Constructor argument theme={null}
from gnani.tts import GnaniTTSClient
client = GnaniTTSClient(api_key="your-api-key")
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.tts import GnaniTTSClient
# Picks up GNANI_API_KEY automatically
client = GnaniTTSClient()
```
### Synthesize Speech
The `synthesize` method returns the complete audio as bytes, which you can write to a file or pass directly to an audio player.
```python theme={null}
from gnani.tts import GnaniTTSClient
client = GnaniTTSClient(api_key="your-api-key")
audio = client.synthesize(
"नमस्ते, आप कैसे हैं?",
voice="sia",
)
with open("output.wav", "wb") as f:
f.write(audio)
```
### Custom Audio Config
Control the sample rate, encoding, and container format of the output audio.
```python theme={null}
from gnani.tts import GnaniTTSClient, AudioConfig
client = GnaniTTSClient(api_key="your-api-key")
audio = client.synthesize(
"यह एक टेस्ट है",
voice="raju",
audio_config=AudioConfig(
sample_rate=44100,
encoding="linear_pcm",
container="wav",
),
)
with open("output.wav", "wb") as f:
f.write(audio)
```
### List Available Voices
```python theme={null}
from gnani.tts import GnaniTTSClient
voices = GnaniTTSClient.supported_voices()
print(voices)
```
## Supported Languages
The Gnani Timbre v2.0 API supports 2 languages.
| Language | Native Script | Example |
| -------- | ------------------- | -------------------------- |
| English | Latin | "I am going to the market" |
| Hindi | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
# Text-to-Speech (Streaming)
Source: https://docs.gnani.ai/vachana/TTS/tts-sse
POST /api/v1/tts/sse
Stream audio in chunks as it's generated via Server-Sent Events.
**Currently in beta.** You're on the priority waitlist and among the first to get access.
## Overview
Receive audio in chunks as it's generated, allowing playback to start immediately. Reduces latency compared to [TTS REST](/vachana/TTS/tts-inference).
Passing numbers, IDs, dates, or currency as raw strings causes mispronunciations. See the [Input Formatting Guide](/vachana/TTS/tts-input-formating) for correct formatting of phone numbers, account numbers, PINs, Aadhaar, vehicle registration numbers, GSTIN, currency, and more.
***
## Available Voices
| Voice | Gender | Description |
| :------ | :----- | :----------------------- |
| Pranav | Male | Bold, Trustworthy |
| Kaveri | Female | Confident, Bright |
| Shubhra | Female | Gentle, Expressive |
| Deepak | Male | Grounded, Conversational |
***
## Python SDK
The SDK's streaming client handles SSE parsing and chunk reassembly for you — you just iterate and write.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
```python Constructor argument theme={null}
from gnani.tts import GnaniTTSStreamClient
client = GnaniTTSStreamClient(api_key="your-api-key")
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.tts import GnaniTTSStreamClient
client = GnaniTTSStreamClient()
```
### Stream Audio to a File
`synthesize_stream` yields audio chunks as they arrive. Playback or writing can begin before the full response is complete.
```python theme={null}
from gnani.tts import GnaniTTSStreamClient
client = GnaniTTSStreamClient(api_key="your-api-key")
with open("output.wav", "wb") as f:
for chunk in client.synthesize_stream(
"Streaming TTS response in Hindi",
voice="sia",
):
f.write(chunk)
```
### With Custom Audio Config
```python theme={null}
from gnani.tts import GnaniTTSStreamClient, AudioConfig
client = GnaniTTSStreamClient(api_key="your-api-key")
with open("output.wav", "wb") as f:
for chunk in client.synthesize_stream(
"नमस्ते, आप कैसे हैं?",
voice="raju",
audio_config=AudioConfig(
sample_rate=44100,
encoding="linear_pcm",
container="wav",
),
):
f.write(chunk)
```
## Supported Languages
The Gnani Timbre v2.0 API supports 2 languages.
| Language | Native Script | Example |
| -------- | ------------------- | -------------------------- |
| English | Latin | "I am going to the market" |
| Hindi | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |
# Text-to-Speech (Realtime)
Source: https://docs.gnani.ai/vachana/TTS/tts-websocket
Real-time text-to-speech with streaming audio via WebSocket.
**Currently in beta.** You're on the priority waitlist and among the first to get access.
## Overview
Stream audio in real-time with the lowest latency. Perfect for interactive assistants and live applications. For simpler use cases, see [TTS REST](/vachana/TTS/tts-inference) or [TTS SSE](/vachana/TTS/tts-sse).
Passing numbers, IDs, dates, or currency as raw strings causes mispronunciations. See the [Input Formatting Guide](/vachana/TTS/tts-input-formating) for correct formatting of phone numbers, account numbers, PINs, Aadhaar, vehicle registration numbers, GSTIN, currency, and more.
## Available Voices
| Voice | Gender | Description |
| :------ | :----- | :----------------------- |
| Pranav | Male | Bold, Trustworthy |
| Kaveri | Female | Confident, Bright |
| Shubhra | Female | Gentle, Expressive |
| Deepak | Male | Grounded, Conversational |
## Endpoint
```text theme={null}
wss://api.vachana.ai/api/v1/tts
```
## Authentication
All Realtime connections require the following headers:
| Header | Required | Description | Example |
| -------------- | -------- | ------------------------------- | ------------------- |
| `Content-Type` | Yes | Must be `application/json` | `application/json` |
| `X-API-Key-ID` | Yes | Your API key for authentication | `` |
## Request Format
Send a JSON message with the following structure:
```json theme={null}
{
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-voice-v3",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm"
}
}
```
Number of audio channels (e.g., `1` for mono, `2` for stereo)
Sample width in bytes (e.g., `2` for 16-bit audio)
Audio encoding format (e.g., `linear_pcm`)
Audio container format (e.g., `wav`)
## Response
The server streams audio data in real-time as binary chunks. Each chunk contains PCM audio data according to the specified `audio_config`.
## Example Usage
```javascript JavaScript theme={null}
const ws = new WebSocket("wss://api.vachana.ai/api/v1/tts", {
headers: {
"Content-Type": "application/json",
"X-API-Key-ID": "",
},
});
ws.on("open", () => {
const request = {
text: "नमस्ते, आप कैसे हैं?",
model: "vachana-voice-v3",
audio_config: {
sample_rate: 44100,
encoding: "linear_pcm",
},
};
ws.send(JSON.stringify(request));
});
ws.on("message", (data) => {
// Handle audio chunks
console.log("Received audio chunk:", data);
});
ws.on("error", (error) => {
console.error("WebSocket error:", error);
});
ws.on("close", () => {
console.log("WebSocket connection closed");
});
```
```python Python theme={null}
import websocket
import json
def on_message(ws, message):
# Handle audio chunks
print(f"Received audio chunk: {len(message)} bytes")
def on_error(ws, error):
print(f"Error: {error}")
def on_close(ws, close_status_code, close_msg):
print("WebSocket connection closed")
def on_open(ws):
request = {
"text": "नमस्ते, आप कैसे हैं?",
"model": "vachana-voice-v3",
"audio_config": {
"sample_rate": 44100,
"encoding": "linear_pcm"
}
}
ws.send(json.dumps(request))
ws = websocket.WebSocketApp(
"wss://api.vachana.ai/api/v1/tts",
header={
"Content-Type": "application/json",
"X-API-Key-ID": ""
},
on_open=on_open,
on_message=on_message,
on_error=on_error,
on_close=on_close
)
ws.run_forever()
```
***
## Python SDK
The SDK's realtime client manages the WebSocket lifecycle, audio streaming, and async iteration so you can focus on your application logic.
### Installation
```bash theme={null}
pip install gnani-vachana
```
Requires **Python 3.9+**.
### Authentication
```python Constructor argument theme={null}
from gnani.tts import GnaniTTSRealtimeClient
client = GnaniTTSRealtimeClient(api_key="your-api-key")
```
```bash Environment variable theme={null}
export GNANI_API_KEY="your-api-key"
```
```python Environment variable (usage) theme={null}
from gnani.tts import GnaniTTSRealtimeClient
client = GnaniTTSRealtimeClient()
```
### Stream Audio Chunks in Real-Time
Use the async context manager to open the connection and iterate over audio chunks as they arrive.
```python theme={null}
import asyncio
from gnani.tts import GnaniTTSRealtimeClient
async def main():
async with GnaniTTSRealtimeClient(api_key="your-api-key") as client:
with open("output.wav", "wb") as f:
async for chunk in client.synthesize(
"नमस्ते, आप कैसे हैं?",
voice="sia",
):
f.write(chunk)
asyncio.run(main())
```
### Collect All Audio at Once
If you don't need to process chunks as they arrive, use `synthesize_and_collect` to get the full audio as a single bytes object.
```python theme={null}
import asyncio
from gnani.tts import GnaniTTSRealtimeClient
async def main():
async with GnaniTTSRealtimeClient(api_key="your-api-key") as client:
audio = await client.synthesize_and_collect(
"Realtime TTS response",
voice="neha",
)
with open("output.wav", "wb") as f:
f.write(audio)
asyncio.run(main())
```
## Supported Languages
The Gnani Timbre v2.0 API supports 2 languages.
| Language | Native Script | Example |
| -------- | ------------------- | -------------------------- |
| English | Latin | "I am going to the market" |
| Hindi | Devanagari (हिन्दी) | "मैं बाज़ार जा रहा हूँ" |