Delivered an
IRC82-Compliant
Dataset for Highway Infrastructure Monitoring AI Model

An Indian Government-Backed Infrastructure Consulting Company
The client is an engineering consultancy with over 25 years of experience across transportation, energy, and urban development. It provides project management, infrastructure design, advisory, and digital transformation support for highways, tunnels, railways, metros, hydropower projects, solar energy projects, and urban infrastructure. Many of its engagements align infrastructure delivery with public-sector mandates and national development programs. The organization is now extending this expertise into data-led infrastructure management by building scalable pipelines for government initiatives focused on AI-based assessment and monitoring.
Training Datasets for a Highway Damage Detection AI Model
As part of a public-sector program involving the NHAI (National Highways Authority of India) and a state road construction agency, the client required image labeling services for road infrastructure. The objective was to prepare standardized training data for an infrastructure monitoring AI model supporting road assessment and maintenance planning.
Survey vehicles captured highway images across several Indian states, which the client supplied to our team in bulk batches. We performed all highway image annotations within the client-hosted and administered CVAT environment.
The project specifications covered:
Labeling Methods
We used bounding boxes and four-point or multi-point polygons to identify, outline, classify, and differentiate road assets and surface damage.
Classification Framework
Apply 71 predefined labels covering defects such as potholes, edge deformation, rutting, and alligator cracking, alongside signage, guardrails, road markings, and drainage assets.
Engineering Standard
Follow IRC82 requirements to maintain consistent classification logic across the resulting infrastructure datasets.
Project Scale
Label highway survey images covering more than 1,000 kilometers of road network across several Indian states.
Maintaining IRC82 Accuracy across a Huge Volume of Highway Images
Unlike standard object labeling, this assignment required annotators to distinguish road defects and infrastructure assets according to IRC82. The resulting data would support government highway assessment and maintenance projects, leaving no room for inconsistent or approximate labeling. Our image labeling services therefore had to maintain the same taxonomy, boundary rules, and classification logic despite changing road conditions and irregular data volumes.
Managing Irregular Batch Volumes across Diverse Highways
Highway images arrived in large batches at irregular intervals, causing sudden workload spikes. Each road corridor also differed in surface composition, visible deterioration, surrounding terrain, and infrastructure assets. These differences made it difficult to standardize throughput without accounting for each dataset's context.
Separating Visually Similar Defects across 71 IRC82 Classes
The labeling framework contained 71 asset and damage classes, including closely related forms of cracking, patching, deformation, potholes, and broken edges. Annotators had to evaluate visual patterns carefully because similar-looking defects could represent different engineering conditions. The challenge extended beyond identifying damage to assigning the correct IRC82 classification.
Maintaining Label Accuracy under Variable Survey Conditions
The survey covered more than 1,000 kilometers across different geographic regions and roadway conditions. Low-light underpasses, glare from wet surfaces, dusty shoulders, changing pavement types, and inconsistent contrast often reduced feature visibility. Maintaining precise boundaries and consistent classification logic across these images was a significant highway image annotation challenge.
Preventing Hidden Annotation Drift during Team Expansion
After the pilot was approved, the expanded workload required additional annotators. The central challenge was ensuring that new team members interpreted IRC82 requirements exactly like the original specialists. An experienced annotator might recognize a long depression along a vehicle wheel path as rutting and trace its complete extent with a multi-point polygon. A newly onboarded annotator could mistake the deepest section for a pothole, label only that localized area, and apply a bounding box. Both outputs might appear complete during a basic review, despite using inconsistent classification and boundary logic. If repeated across batches, this hidden labeling drift could compromise the reliability of the entire infrastructure dataset.
A Quality-First Framework for Large-Scale Road Infrastructure Annotation
To scale image labeling services for road infrastructure without introducing inconsistent interpretations, we established an approved baseline before expanding capacity. The initial pilot included four experienced annotators and one quality reviewer with prior exposure to infrastructure training data. The pilot let us test the complete workflow within the client-managed CVAT environment. Its output also became the reference dataset for evaluating production batches after the scale-up. Once the pilot demonstrated the required quality, we expanded the operation to 35 full-time annotators and seven quality reviewers. The complete team comprised 42 specialists.
Training before Production
Every annotator completed a compulsory training and calibration program before working on live highway images. The program covered all 71 IRC82 damage and asset categories through detailed guidelines and worked examples. Annotators completed targeted exercises on categories with similar visual characteristics. These included alligator and longitudinal cracking, as well as potholes and rutting. They also practiced on sample images from the client’s own dataset. Calibration sessions with the client’s project team clarified assignment-specific conventions. This ensured that the expanded team followed the same interpretation established during the pilot.
Applying Rules for Maintaining Image Quality
We introduced clear decision rules for three image-quality levels: clear, degraded, and unworkable. Clear images contained sufficient visual information and entered the normal annotation workflow. Degraded images contained issues such as poor lighting, glare, low contrast, or dusty shoulders. Annotators processed identifiable features but flagged genuinely uncertain cases for review instead of assigning assumptions-based labels. Unworkable images did not provide enough visual evidence to identify road surface features reliably. We escalated these cases to the client for a final decision. The client could request another survey of the affected segment or exclude the images from the infrastructure dataset.
Operating within Client-Managed CVAT Instance
All labeling remained inside the client’s hosted and administered CVAT environment. The client retained oversight of user access, annotation versions, and export pipelines throughout production. This supported the data governance requirements associated with government infrastructure projects. Our team processed more than 350,000 highway images using bounding boxes and four-point or multi-point polygon annotation. Approximately 80% of annotations involved uniform assets suited to four-point boxes. The remaining 20% involved irregular road defects requiring detailed multi-point tracing. This allocation maintained pixel-level accuracy where complex boundaries mattered while protecting throughput and project timelines.
Reviewing Every Batch
The quality framework combined annotator verification with specialist-led review. Each annotator checked completed batches before submission and identified uncertain annotations for peer assessment. Quality reviewers with civil engineering backgrounds then audited structured samples from every batch. Reviews measured the output against both the IRC82 taxonomy and the pilot reference dataset. We recorded any differences in classification or boundary placement and fed them back to the team through calibration sessions. This continuous feedback prevented isolated errors from becoming repeated labeling patterns.
Expanding into OPRMC Monitoring
After we completed the original highway assignment, the client expanded our role to the OPRMC (Output and Performance-Based Road Maintenance Contract) initiative. This contract model evaluates contractors according to actual road conditions and performance outcomes. This differs from conventional contracts that calculate payments from material quantities or completed work volumes. Under OPRMC, the assessment software needed to verify contractor performance against defined KPIs. We labeled more than 300,000 additional highway images (over 1 million annotations) to identify defects and signs of deterioration. This extended scope produced over one million annotations for automated road performance assessment.
Project Outcomes
Completed 3M+ Annotations
We produced approximately two million annotations for national highway corridor assessment and another one million for OPRMC road condition monitoring and contractor performance evaluation.
Maintained 99% Annotation Accuracy
Accuracy remained consistent throughout the engagement, including the expansion from the five-member pilot team to a 42-specialist production team.
Labeled 650,000+ Highway Images
We completed the image annotation work without any decline in labeling quality, classification consistency, or adherence to IRC82 requirements.
Turn Complex Infrastructure Imagery into Model-Ready Datasets
Building an infrastructure monitoring AI model? Our image annotation services cover highway image annotation, geospatial data annotation, satellite image annotation, and aerial image labeling. Share your requirements to receive a sample aligned with your taxonomy and quality targets.
