Multilingual financial NLP for a chatbot fallback problem
A financial institution's assistant fell back to humans too often on non-English and mixed-language queries.
The challenge and pain points
Intent and entity coverage was weak across the languages its customers actually used.
- The assistant handed off to human agents too often on non-English and mixed-language queries, which raised support cost.
- Intent and entity coverage was weak in the languages customers actually used, so the model misread common requests.
- Financial terms and entities were annotated by generalists without domain review, so labels missed sector-specific meaning.
- Mixed-language queries, where customers switch between English and a local language, fell outside the training data entirely.
- Every unnecessary fallback pushed a routine query to a human, which did not scale with customer growth.
What Corpshore did
Native annotators built intent, NER, and preference data across the target languages with domain review.
- 1Audit the fallback queries
Reviewed where the assistant was falling back and grouped the failures by language and intent to target the real gaps.
- 2Build intent and NER coverage
Native annotators built intent classification and named-entity data across the target languages, covering the requests customers actually make.
- 3Add domain review
Applied financial-domain review so entities and intents carried sector-specific meaning rather than a generic reading.
- 4Produce preference data for responses
Built RLHF preference data so the assistant learned which multilingual responses customers judged as resolving the query.
- 5Gate through the QA cascade
Ran every batch through the three-tier QA cascade with native and domain reviewers before delivery.
The solution
Corpshore built the missing coverage in the languages the assistant was failing on, with native annotators producing intent and named-entity data for the requests customers actually make. Financial-domain review made sure entities carried sector-specific meaning rather than a generic one.
Mixed-language queries were treated as first-class rather than edge cases, so the assistant learned to handle the code-switching customers use in practice. RLHF preference data taught the model which multilingual responses customers judged as resolving their query.
Every batch passed the three-tier QA cascade with native and domain reviewers. The assistant's fallback rate on multilingual queries came down and resolution improved, though the exact fallback numbers are confidential to the client.
Results
The assistant fell back to human agents less often on multilingual and mixed-language queries, and resolution on those queries improved, once native intent, NER, and preference data were in place. The precise fallback rate is confidential to the client, so the outcome is described qualitatively here. The data was built by native, in-region annotators at Corpshore's typical cost advantage versus US-domestic providers.
| Metric | Result | Notes |
|---|---|---|
| Language coverage | Multilingual | Built for the languages customers use |
| Chatbot fallback | Reduced | Client figure confidential; qualitative |
| Multilingual resolution | Improved | Client figure confidential; qualitative |
| Data types | NER + intent + RLHF | Domain-reviewed |
| Annotators | Native, in-region | With financial-domain review |
| QA | Three-tier cascade | Native and domain reviewers |
| Cost vs US-domestic | 50 to 70% lower | Methodology-level figure, not a client quote |
Illustrative cost index, Corpshore set to 100. The band reflects Corpshore's typical 50 to 70% cost advantage versus US-domestic providers. Representative of methodology, not exact client data.
Representative distribution of caught errors across the cascade, locking in 97%+ accuracy before delivery.
Client names and some figures are confidential. Where an exact client metric is not published, outcomes are described qualitatively and charts show representative or methodology-level data, including the three-tier QA cascade distribution and Corpshore's typical 50 to 70% cost advantage versus US-domestic providers.
Ready to scope a pilot?
Tell us your modality, volume, and languages. We'll return an indicative scope, timeline, and cost band.