Role Description
This role is for one of our clients.
Compensation: $12.68 per hour
Fluent Language Skills Required: Odiya. Native fluency in Odiya (Oriya), including full command of Odiya script and orthography, is required for this position. All annotation and transcription work is performed in Odiya.
Why This Role Exists:
-
Document understanding breaks down fastest in languages that parsing and vision-language models rarely see.
-
This project builds training data for Odiya, alongside four other Indic scripts, Japanese, and Korean.
-
Each task takes a real, publicly available PDF page and produces a complete structural map of that page, paired with a faithful transcription of every text region in the original script.
-
The dataset deliberately concentrates on material models handle worst: handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages.
-
Documents are drawn from various sources to reflect the real diversity of Odiya documents.
-
Delivered work is human-authored throughout, with all component identification, typing, reading order, and transcription performed by people.
What You'll Do:
-
Open and check a task: confirm the page is in Odiya, is legible, has real content, and shows no personal details.
-
Annotate structure: identify and bound every meaningful region of the page and assign each a component type and a reading-order index.
-
Record relationships: link each region to the figure or table it belongs to through a parent component identifier.
-
Transcribe faithfully: reproduce all text exactly as it appears in Odiya script, including handwritten content.
-
Capture page metadata: language, document type, source, page dimensions, and flags for tables, formulas, and handwriting.
-
Review a colleague's work: every task is reviewed end to end by a second Odiya expert.
Qualifications
-
You are a native Odiya speaker with full command of the script, its diacritics, and its conjunct forms.
-
You have worked with documents: annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
-
You are exact: character-level accuracy matters more here than speed.
-
You are systematic: you apply a taxonomy consistently across hundreds of pages.
-
You are comfortable with unfamiliar layouts: multi-column newspapers, exam papers, handwritten forms.
Nice-to-Have Specialties
-
Regional-language AI data: annotation, labeling, grading, or bilingual evaluation for training datasets.
-
Transcription and localization: MTPE, subtitling, bilingual QA, OCR correction, or post-editing.
-
Document production: typesetting, copy-editing, proofreading, or digitization of Odiya-language material.
-
Script and encoding: Unicode normalization, Odiya input methods, numeral-form and character-form accuracy.
What Success Looks Like
-
Every meaningful region on the page is captured, correctly bounded, and correctly typed.
-
Reading order reflects how the page is actually read, including across columns.
-
Transcriptions match the source character for character, in Odiya script rather than transliteration.
-
Your tasks pass second-expert review the first time.
-
Unsuitable pages are flagged up front rather than after thirty minutes of work.