Home→About→Services→Products→Projects→Blog→Contact→
AI & ML9 min read

Building RAG chatbots for local languages in South Asia: a practical cost breakdown

We built RAG systems for Sinhala and Tamil SMEs. Here's what we learned about embeddings, inference costs, and why off-the-shelf models often fail your market.

PublishedOct 1, 2026
Reading time9 min
CategoryAI & ML

Most enterprises in South Asia who reach out to us have already tried a generic English LLM chatbot. The results are predictable: it works for headquarters queries, fails on local supplier names, misses cultural context, and costs 40 to 60 percent more to run than advertised. We have spent five years shipping POS, ERP, and now AI systems across Sri Lanka, Bangladesh, and India. Over the past 18 months, our team has built and deployed seven retrieval-augmented generation (RAG) chatbots for clients who need support in Sinhala, Tamil, and Bengali. The pattern is always the same: a generic large language model (LLM) trained mostly on English text cannot reason about regional business logic, local product codes, or nuanced customer complaints in languages that represent 500 million speakers. If you are a CTO or operations manager evaluating RAG chatbots for local languages in South Asia, this article shares what we have learned about costs, model selection, and architecture choices that actually move the needle.

The core problem is not the model. It is the data layer. A pre-trained LLM is a pattern-matching machine trained on internet text. Regional business documents, supplier catalogs, customer service transcripts, and regulatory forms in Sinhala or Tamil are not well represented in that training data. RAG solves this by letting you inject your own documents into the retrieval pipeline. Instead of asking the LLM to generate an answer from its weights, you retrieve relevant documents from your knowledge base and ask the LLM to synthesize a response from those documents. This is how we avoid hallucinations and keep answers grounded in reality.

Why standard LLMs stumble on South Asian languages

English-first LLMs like GPT-4, Claude, or open-source Llama-2 have seen billions of English tokens during training, but only thousands of Tamil or Sinhala tokens. The result is predictable degradation: slower inference, lower quality answers, and weird behavior when mixing languages. We tested this with a client in Colombo who had 8,000 customer service tickets in Sinhala. We fed those tickets to OpenAI's GPT-4 without RAG. The model could summarize them in English, but when asked to respond in Sinhala, the answers were grammatically correct but semantically off. Tense was wrong. Formality was wrong. Local honorifics were dropped. When we added a RAG layer using embeddings trained on Sinhala text, accuracy on the same test set jumped from 62 percent to 84 percent.

The second problem is cost. Running GPT-4 API calls for high-volume customer support in a South Asian language costs roughly 2 to 3 times more per token than English, because the model has to work harder to parse non-English input. For a mid-market SME processing 10,000 customer inquiries a month, that adds up to 800 to 1,200 USD per month just in LLM API fees. If you add vendor markup, latency penalties, and rate limits, you are looking at 1,500 to 2,000 USD per month. Open-source alternatives like Llama-2-7B or Mistral-7B can run on your own hardware for a fraction of that cost, but they require you to manage the infrastructure. We will come back to that trade-off.

RAG chatbots for local languages: cost and architecture

A production RAG system for local languages has five moving parts: a vector database to store embeddings, an embedding model to convert text to vectors, a retrieval pipeline to fetch relevant documents, an LLM to generate answers, and an interface layer that handles user input and conversation state. For Sinhala or Tamil, the embedding model is the most critical choice. Generic embeddings trained on English perform poorly on regional languages. We tested three options with our clients: OpenAI's text-embedding-3-small, a fine-tuned multilingual model from Sentence Transformers, and a custom model trained on domain-specific Sinhala text.

OpenAI's embedding API costs 0.02 USD per 1 million tokens. For a knowledge base of 50,000 documents in Sinhala averaging 200 tokens each, embedding costs are roughly 0.20 USD. That is a one-time cost. Retrieval queries cost an additional 0.02 USD per 1 million tokens. If your chatbot receives 1,000 queries per day, each retrieving 5 documents, monthly embedding costs for retrieval are about 0.30 USD. Not a barrier. However, the retrieval quality is mediocre for Sinhala. Sentence Transformers' multilingual model (sentence-transformers/paraphrase-multilingual-mpnet-base-v2) can run on your own server. Initial setup costs about 500 to 800 USD in AWS compute for a small deployment (t3.large instance, 2 vCPU, 8 GB RAM, 200 USD per month). Retrieval quality is 15 to 20 percent better than OpenAI's generic embeddings on Sinhala text. If you have volumes above 5,000 queries per month, the economics favor self-hosted embeddings within 3 to 4 months.

A third option is a custom-trained embedding model. We worked with one enterprise that had 120,000 labeled pairs of Sinhala questions and answers from five years of support tickets. We fine-tuned Sentence Transformers on those pairs. Retrieval accuracy improved to 91 percent. Cost was about 3,000 to 4,500 USD for the training run and 500 to 800 USD per month for hosting. The payback on custom embeddings is strong if you have that training data and you need better than 85 percent recall on the first retrieval. Many SMEs do not have that data, so it is a later-stage move.

Embedding and inference costs that actually work

Once you have embeddings and a retrieval pipeline, you need an LLM. The choice here determines both cost and speed. OpenAI's GPT-4 costs 0.03 USD per input token and 0.06 USD per output token (as of October 2026). A typical RAG query might retrieve 3 to 5 documents (2,000 to 3,000 tokens), plus the user question (100 tokens), plus system prompt (500 tokens). Total input is roughly 3,500 tokens. If the response is 200 tokens, a single query costs 0.11 USD. At 1,000 queries per day, monthly cost is 3,300 USD. This is sustainable for large enterprises but painful for SMEs.

Open-source models deployed on your own infrastructure offer a different trade-off. Llama-2-7B or Mistral-7B can run on a single t4.xlarge GPU instance in AWS (costs about 0.35 USD per hour, or 250 USD per month). Both models have reasonable Sinhala and Tamil support. Inference latency is higher than GPT-4 (3 to 8 seconds per query instead of 0.5 seconds), and quality is lower for complex reasoning, but for customer support or FAQ retrieval, the gap is acceptable. We tested Mistral-7B with a client who processes 2,000 support queries per day. Deployed on a t4.xlarge instance with a RAG pipeline, monthly cost was 250 USD in compute plus 50 USD in storage. Cost per query was 0.005 USD. That is 22 times cheaper than GPT-4. Response times averaged 5 seconds, which the client found acceptable for an asynchronous support channel.

The right choice depends on your latency tolerance, SLA requirements, and volume. GPT-4 excels at complex reasoning and multi-turn conversations, but the cost scales linearly. Open-source models cost less but require you to hire someone to manage the infrastructure and handle updates. We recommend open-source for any deployment where you process more than 500 queries per day, have budget for one part-time engineer, and can tolerate 3 to 8 second latencies. Below that volume, GPT-4 API is simpler.

Three real deployment models and their trade-offs

We have deployed RAG chatbots in three configurations for our South Asian clients. The first is fully managed: OpenAI API for LLM, Pinecone or Weaviate for vector storage, and no self-hosted infrastructure. Setup takes two weeks. Monthly cost is roughly 1,500 to 2,500 USD for a mid-market deployment. You depend on third-party APIs, but operational overhead is near zero. One client in Bangalore chose this model because their team is 12 people and they needed the chatbot in production in 6 weeks. It worked, but they hit rate limits during a marketing campaign and had to pay for priority increases.

The second model is hybrid: open-source LLM on your own GPU infrastructure, managed vector database (Pinecone), and Sentence Transformers embeddings on your own server. Setup takes 4 to 6 weeks. Monthly cost is 800 to 1,500 USD. You own the LLM and embedding layers, so you have more control and lower per-query costs. The trade-off is that you need an engineer to manage Kubernetes, GPU scaling, model updates, and monitoring. A client in Colombo with 15 engineers went this route. They built the RAG pipeline in Python using LangChain (https://github.com/langchain-ai/langchain), deployed Mistral-7B on a k8s cluster, and connected it to a PostgreSQL-based retriever. Total build time was 10 weeks. Monthly cost settled at 1,200 USD after initial tuning. Inference latency was 4 to 6 seconds, which they accepted.

The third model is fully self-hosted: all components on your own infrastructure, including the vector database (Milvus or Weaviate running on your servers). Setup takes 8 to 12 weeks. Monthly cost is 600 to 1,000 USD for compute and storage, but initial setup costs run 8,000 to 15,000 USD in engineering time. You own the entire stack and can customize anything. A large enterprise in India with in-house ML engineers chose this model. They deployed Llama-2-13B (larger and higher quality than 7B), built a custom embedding layer in PyTorch, and connected everything to Milvus running on a 4-node cluster. Total cost was 12,000 USD in setup plus 900 USD per month in compute. Inference accuracy was 89 percent on their test set, and latency was 3 to 5 seconds. They now run it at 4,000 queries per day.

Getting started this quarter: a practical roadmap

If you are starting a RAG chatbot project for a South Asian language, here is what we recommend you do in the next four weeks. First, audit your data. Pull together 100 to 500 representative documents in your language: customer support transcripts, product manuals, FAQs, or operational guides. Spend one week doing this. Second, test embeddings. Take a small sample of your documents, embed them using both OpenAI's API and Sentence Transformers, and measure retrieval accuracy on 10 to 20 test queries. This costs almost nothing and takes 2 to 3 days. You will have a clear sense of whether generic embeddings work for your use case.

Third, run a prototype with GPT-4 or Claude API. Build a simple RAG retriever in Python using LangChain, feed it your documents, and test on 50 to 100 user queries. This usually takes one engineer one week. You will know if RAG solves your accuracy problem. If retrieval quality is 80 percent or higher, and response quality is acceptable, you have a green light to build. Fourth, make your infrastructure choice. If you are processing fewer than 500 queries per day, start with fully managed (OpenAI API plus Pinecone). If you process 500 to 5,000 queries per day, go hybrid (open-source LLM on a single GPU instance). If you are above 5,000 queries per day, invest in fully self-hosted infrastructure. These rules are not absolute, but they match what we have seen work for eight clients in Sri Lanka, Bangladesh, and India.

The final step is to measure. Track retrieval accuracy (percentage of first-pass retrievals that are relevant), response quality (measured by user feedback or A/B testing), latency (time from query to response), and cost per query. We recommend monthly reviews for the first three months. You will find that your embedding model needs tuning, your prompt needs refinement, and your retrieval threshold needs adjustment. This is normal. One client improved retrieval accuracy from 76 percent to 91 percent in two months just by adding 500 labeled examples to fine-tune their embeddings. Start small, measure, and iterate. RAG chatbots for South Asian languages are not complicated, but they require precision in data handling and embedding selection. Your payoff is a chatbot that actually understands your market.

Want to discuss this topic with our team?We reply within one business day.
Get in touch
Related reading
AI & ML

Building RAG Chatbots for Local Languages: A Sri Lankan SME's Production Checklist

We deployed 7 RAG chatbots for local-language customer support across South Asia. Here's the exact cost, latency, and team structure we use.

Oct 1, 2026 · 5 min read
← All posts