CapsAI
Back to tools
Research data pending

Indian English Subtitle Accuracy Benchmark

This is a methodology shell, not a published benchmark. No accuracy scores, rankings, or comparative claims are available yet.

Version 0.1Prepared 19 August 2026Author and editor: CapsAI research team

Proposed scope

What a responsible study must define

Representative speech

Document speaker regions, recording conditions, content domains, code-switching, and consent rather than presenting one accent as all Indian English.

Transparent ground truth

Use trained human transcription and an adjudication process for names, numbers, punctuation, and ambiguous speech.

Multiple useful measures

Report word error rate alongside proper-noun, numeric, segmentation, and subtitle-readability observations.

Reproducible reporting

Publish sample definitions, exclusions, tool versions, evaluation scripts, uncertainty, and limitations with the results.

Dataset design

Accent, language, and recording categories

The final study must publish its sampling rationale and cannot claim to represent all Indian English from a narrow speaker set.

Regional coverage

Multiple regions and documented speaker backgrounds

Code-switching

English-only and English mixed with Indian languages

Content domains

Creator, education, business, interview, and instructional speech

Audio conditions

Studio, phone, room noise, music, compression, and overlapping speech

Test methodology

Proposed evaluation sequence

  1. 1Create consented, documented evaluation audio that is not part of system tuning.
  2. 2Produce human reference transcripts with independent review and adjudication.
  3. 3Run every compared system with disclosed settings and version information.
  4. 4Normalize only pre-declared formatting differences before scoring.
  5. 5Publish aggregate and category-level observations with uncertainty and limitations.

Scoring framework

More than one headline metric

Word error rate

Standard insertion, deletion, and substitution measurement after documented normalization.

Proper-noun accuracy

Separate review of names, brands, products, and places.

Number accuracy

Dates, currency, percentages, measurements, and other meaning-sensitive figures.

Subtitle segmentation

Readable cue boundaries, line breaks, duration, overlap, and reading speed.

Code-switch handling

Errors around transitions between Indian English and other languages.

Sample output component

How evaluated examples will be disclosed

Reference transcript

Pending human-adjudicated text

System output

Pending evaluated system text

Error annotation

Pending substitution, deletion, insertion, name, number, and segmentation labels

Placeholder comparison table with no benchmark results
SystemVersionOverall WERNamesNumbersStatus
CapsAIPending---Data pending
Comparison system APending---Data pending
Comparison system BPending---Data pending

Methodology disclosure

The eventual report must disclose dataset construction, consent, exclusions, human-reference process, normalization, scoring software, system versions, conflicts of interest, uncertainty, limitations, and any data unavailable for independent inspection. CapsAI evaluating its own system is a conflict that must remain visible.

Reference standards

Publication gate

This page will remain noindex and outside the XML sitemap until a completed study includes evidence, methodology, limitations, review, and results that support every public claim.

Start with 3 free minutes

Test CapsAI with Your Own Video

Upload your own Indian English video and try CapsAI's subtitle workflow. Start with 3 free minutes.