Skip to content
A Top 5 global AI outsourcing company by Outsource Accelerator, above Scale AI.
Fintech · Multi-hub · Text / NER + RLHF

Multilingual financial NLP for a chatbot fallback problem

A financial institution's assistant fell back to humans too often on non-English and mixed-language queries.

Multilingual
Coverage
Reduced
Fallback
NER + RLHF
Data types
Multi-hub
Region

The challenge and pain points

Intent and entity coverage was weak across the languages its customers actually used.

  • The assistant handed off to human agents too often on non-English and mixed-language queries, which raised support cost.
  • Intent and entity coverage was weak in the languages customers actually used, so the model misread common requests.
  • Financial terms and entities were annotated by generalists without domain review, so labels missed sector-specific meaning.
  • Mixed-language queries, where customers switch between English and a local language, fell outside the training data entirely.
  • Every unnecessary fallback pushed a routine query to a human, which did not scale with customer growth.

What Corpshore did

Native annotators built intent, NER, and preference data across the target languages with domain review.

  1. 1
    Audit the fallback queries

    Reviewed where the assistant was falling back and grouped the failures by language and intent to target the real gaps.

  2. 2
    Build intent and NER coverage

    Native annotators built intent classification and named-entity data across the target languages, covering the requests customers actually make.

  3. 3
    Add domain review

    Applied financial-domain review so entities and intents carried sector-specific meaning rather than a generic reading.

  4. 4
    Produce preference data for responses

    Built RLHF preference data so the assistant learned which multilingual responses customers judged as resolving the query.

  5. 5
    Gate through the QA cascade

    Ran every batch through the three-tier QA cascade with native and domain reviewers before delivery.

The solution

Corpshore built the missing coverage in the languages the assistant was failing on, with native annotators producing intent and named-entity data for the requests customers actually make. Financial-domain review made sure entities carried sector-specific meaning rather than a generic one.

Mixed-language queries were treated as first-class rather than edge cases, so the assistant learned to handle the code-switching customers use in practice. RLHF preference data taught the model which multilingual responses customers judged as resolving their query.

Every batch passed the three-tier QA cascade with native and domain reviewers. The assistant's fallback rate on multilingual queries came down and resolution improved, though the exact fallback numbers are confidential to the client.

Results

The assistant fell back to human agents less often on multilingual and mixed-language queries, and resolution on those queries improved, once native intent, NER, and preference data were in place. The precise fallback rate is confidential to the client, so the outcome is described qualitatively here. The data was built by native, in-region annotators at Corpshore's typical cost advantage versus US-domestic providers.

MetricResultNotes
Language coverageMultilingualBuilt for the languages customers use
Chatbot fallbackReducedClient figure confidential; qualitative
Multilingual resolutionImprovedClient figure confidential; qualitative
Data typesNER + intent + RLHFDomain-reviewed
AnnotatorsNative, in-regionWith financial-domain review
QAThree-tier cascadeNative and domain reviewers
Cost vs US-domestic50 to 70% lowerMethodology-level figure, not a client quote
Labeling cost, Corpshore vs US-domestic (index)
Corpshore100
US-domestic (typical range)≈250 to 333

Illustrative cost index, Corpshore set to 100. The band reflects Corpshore's typical 50 to 70% cost advantage versus US-domestic providers. Representative of methodology, not exact client data.

Where errors get caught (three-tier QA cascade)
Tier 1 · annotator + peer~80%
Tier 2 · expert QA lead~15%
Tier 3 · programmatic + consensusremainder

Representative distribution of caught errors across the cascade, locking in 97%+ accuracy before delivery.

Client names and some figures are confidential. Where an exact client metric is not published, outcomes are described qualitatively and charts show representative or methodology-level data, including the three-tier QA cascade distribution and Corpshore's typical 50 to 70% cost advantage versus US-domestic providers.

Ready to scope a pilot?

Tell us your modality, volume, and languages. We'll return an indicative scope, timeline, and cost band.

Start a pilot Explore careers