All posts
data annotationdata taggingcomputer vision

Annotation Pipelines That Scale: India’s Data-Tagging Edge

Dr Ishit Karoli
March 19, 2026
4 min read· 10 sections
Annotation Pipelines That Scale: India’s Data-Tagging Edge

"AI just needs more labelled data" is the most boring true thing in the field. The interesting question is how you produce that labelled data without quality collapsing under volume. India still has a structural advantage here, but only for teams that run annotation like a real operation, not a body-shop.

Pre-labelling is the throughput multiplier

The biggest lever on annotator throughput is usually pre-labelling. Run your existing model, or a general-purpose one such as SAM 2, Florence-2 or Grounding DINO, on raw items first; humans then correct rather than label from scratch. Throughput typically rises substantially, and quality often improves because annotators review edges rather than place every box.

Watch for automation bias: annotators tend to accept pre-labels that look plausible. Counter it by seeding a few deliberately wrong pre-labels into calibration batches and tracking how many get caught.

Inter-annotator agreement is the quality metric that matters

If your annotators don't agree with each other above a defined threshold, you don't have annotation quality; you have annotation noise. Calibrate every new annotator against a gold set until their agreement clears the threshold you have set for the task, then sample-validate every batch. Track the metric weekly and treat declines like incidents.

Choosing the measure

  • Classification. Cohen's kappa for two annotators; Fleiss' kappa or Krippendorff's alpha for more. They correct for agreement by chance, which raw percentage agreement does not.
  • Boxes and masks. Intersection over union (IoU) against the gold label, with a pass threshold per class.
  • Named entities and spans. Span-level F1 against gold, with exact and partial matches reported separately.
  • Preference labelling. Agreement on pairwise choices, plus how often annotators pick the same side as a known-good reference.

Set thresholds with the model team, based on what the downstream model can tolerate, not on what looks good in a weekly report.

The error taxonomy you need to define on day one

Vague guidelines produce vague labels. Pick a closed taxonomy of error types early (boundary error, class confusion, missed object, hallucinated object) and use it for QA review. The taxonomy itself becomes the training material for new annotators, and error counts by type tell you whether to fix the guideline, the tool or the training.

Edge cases drive guideline iterations

The first version of any annotation guideline is wrong. Edge cases that come up in calibration go into the guideline gallery, with examples and the canonical decision. After a few iteration cycles the guideline stabilises and onboarding gets much faster. Keep a changelog, and when a rule changes, decide explicitly whether past labels need rework.

Tooling: Label Studio for most cases, custom for the rest

Label Studio covers most annotation use cases out of the box: bounding boxes, segmentation, NER and classification. CVAT is a strong pick for video, with interpolation between keyframes. Custom tools are warranted only for unusual modalities such as LiDAR, 3D or multi-modal preference labelling. Don't over-engineer.

Modalities we run regularly

  • Image. Boxes, polygons, segmentation, keypoints. The workhorse.
  • Video. Tracking and action recognition. Considerably slower and harder per item than image, so plan capacity accordingly.
  • Text. NER, classification, RLHF preference labelling. It scales differently: quality is the bottleneck, not throughput.
  • Audio. Speaker diarisation, transcription, sentiment. Multilingual work is where this gets interesting in India.

Running it as an operation

What separates a scalable annotation operation from a body-shop is structure:

  • Tiers. Annotators, reviewers and a QA lead, with reviewers drawn from the best-calibrated annotators.
  • Visible metrics. Throughput and quality tracked per person, shared with them and used for coaching.
  • Secure environment. No local downloads, access by project, and an audit of who saw which items. This is essential for medical, financial or customer data.
  • Batch contracts. Each batch has a definition of done, an acceptance sample and an agreed turnaround.
  • Feedback loop. Model errors in production are traced back to label issues and fed into the guideline.

Where India's advantage is real

The advantage is a combination of a large English-speaking graduate workforce, mature operations management from decades of BPO work, working hours that overlap with Europe and the Middle East, and shift patterns that can cover US mornings. Multilingual audio and text in Indian languages is a particular strength. None of it matters without the discipline above: the advantage is operational, not only cost.

FAQ

How big should the gold set be?

Big enough to cover every class and every known edge case several times. For many projects, a few hundred items is a practical starting point, grown as new edge cases appear.

Can a model label the data instead of people?

For easy, high-confidence items, model labels with sampled human review can work well. For edge cases and new classes, people remain the source of truth.

How we run this at Velura Labs

Our AI Data Tagging service runs dedicated annotation teams out of Lucknow with the discipline above as the default: pre-labelling, calibration, agreement metrics and edge-case galleries. For downstream model training, pair it with our Model Fine-tuning service. Our fine-tuning guide explains when better data outperforms a better model. Talk to us if your model performance is plateauing and you suspect data is the issue.

Velura Labs delivers this for teams across the United States — Seattle (Washington), San Francisco and Los Angeles (California), Austin and Dallas (Texas), and New York — as well as Europe (Paris, Milan, Rome and the wider EU), the Middle East (Dubai, Abu Dhabi and Riyadh) and India. Talk to us wherever you operate.

Now booking Q4 2026

Let's build the
next chapter of your business.

Quick chat on WhatsApp. We'll scope your web, app, or AI build, show you a reference architecture, and price the first slice.

80+
shipped projects
12
industries
ISO 9001:2015
certified
98.4%
CSAT