Doctors in the Loop

Medical Data Annotation Vendor: A Buyer’s Checklist

Verifying medical annotation vendor credentials is the step most vendor evaluation guides skip entirely. Generic “how to vet a data annotation vendor” guides exist, and some are decent: they’ll tell you to ask about QA process, pricing structure, and security posture.

What’s almost constantly missing, because most are written by medical professionals without annotation background or experience, is whether the person labeling your data has the right specialty, whether disagreement between annotators is even measured, and whether any of it would survive a regulator asking to see it.

That is the gap that the HITL project checklist fills, upon going through the scope and a generic annotation guideline. This is what to ask on top of that, specific to medical AI data. We’ll also show you how we’d answer each question ourselves, so you can see what a real answer looks like versus a vague one.

Medical Data Annotation Vendor A Buyer’s Checklist

5 Things to Check When Choosing a Medical Data Annotation Vendor

Before you go deep with any data‑annotation provider, consider asking these five questions. If a vendor can’t answer all five clearly and specifically, it’s a strong signal they may not meet the standards required for a reliable partnership.

  1. Who performs the annotation, and are their credentials and specialties relevant to your data?
  2. How does the vendor measure annotation quality and inter-annotator agreement?
  3. What happens when annotators disagree on a label?
  4. Can the vendor provide documentation and traceability for regulatory or audit purposes? Can you produce documentation that would satisfy an EU AI Act data governance?
  5. How are security, de-identification, and data access handled?

1. Verify clinical expertise and specialty matching

Almost every vendor in the medical AI space uses some version of “medical professionals” or “clinical experts” in their marketing. Very few specify:

  • Are annotators licensed, board-certified, or in training (residents/students)?
  • Are annotators matched to the relevant specialty for your data. For example, a radiologist for imaging, a dermatologist for skin lesion data, a cardiologist for ECG, or is one general “medical annotator” pool used across all modalities?
  • Can the vendor share (even anonymized) confirmation of the specific annotators who would work on your project?

Here’s a worked illustration of why this point materially affects outcomes. (note that this is a composite example built to show the mechanism, not a specific client engagement.)

Two annotators review the same chest X-ray. A general practitioner labels it “no acute findings.” A radiologist flags a subtle early-stage nodule in the periphery of the lung field, the kind of finding that’s easy to miss without pattern recognition built from reading thousands of chest films specifically.

Both annotators are genuinely qualified clinicians. Only one of them has the subspecialty exposure to catch that particular finding reliably. If your model needs to catch rare or early-stage pathology, a generalist-labeled dataset will be systematically blind to the exact cases you most need it to catch, because the error isn’t random, it’s specialty-shaped.

We’d go further than most vendors are willing to: specialty match matters more than credential alone. “Board-certified physician” tells you almost nothing if that physician’s background is family medicine and your dataset is dermoscopic melanoma images. Push past “are they a doctor” to “is this the right kind of doctor for this data,” and don’t accept a vendor’s discomfort with that question as a reason to drop it.

2. Evaluate the vendor's quality assurance methodology

A vendor claiming to deliver a “99% accuracy” on annotation output is meaningless on its own. You need to understand how that number was measured, because accuracy claims at that almost perfect level are rarely achievable in real‑world annotation.

  • What’s your inter-annotator agreement rate (Cohen’s kappa, Fleiss’ kappa, or an equivalent metric) on comparable projects?
  • Is agreement measured against a gold-standard adjudicated label, or just annotator-to-annotator?
  • How many annotators label each item: one, two, or a consensus panel?
  • What’s the adjudication workflow when annotators disagree,  escalation to a senior specialist, majority vote, or discussion consensus?

A vendor with a mature QA process will have this data ready and will be specific about methodology, not just a headline number. A vendor who can only offer “we’re very accurate” hasn’t built the measurement infrastructure to know if that’s true. Read our medical image annotation full guide for AI teams (2026).

3. Documentation and audit-readiness

This is where most annotation vendors fall short, and it’s increasingly where deals are won or lost. Under the EU AI Act, high-risk AI systems will eventually need to meet Article 10 data governance obligations: dataset quality, bias examination, and provenance documentation among them.

Note: that the compliance timeline moved. The “Digital Omnibus on AI,” Regulation (EU) 2026/1744, entered into force July 27, 2026, deferring stand-alone high-risk systems to December 2, 2027 and AI embedded in already-regulated products, including medical devices under MDR/IVDR – to August 2, 2028.  Read what your AI team should know about The EU AI Act’s high-risk requirements .

Even with more runway than teams expected a few months ago, FDA submissions already expect traceable, representative training data under Good Machine Learning Practice, that expectation hasn’t moved, and it’s a more immediate reason to have this conversation with a vendor now rather than closer to either deadline.

Questions to ask:

  • Can you produce a data provenance trail, from raw data intake to final label, for every item in the dataset?
  • Do you document the specific credentials of the annotator(s) who labeled each batch?
  • Have you supported a client through an FDA submission or an EU AI Act conformity assessment before, and can you describe what documentation you provided?
  • Is your QA and adjudication process itself documented in a form that could be handed to a Notified Body or auditor, or does it live only in internal notes?

If your product is heading toward any regulated pathway, this single category of questions probably matters more than price. Remember, annotation you can’t document is annotation you can’t defend later and rework after a failed audit costs far more than the annotation did the first time.

4. Data security and de-identification

Standard, but worth confirming specifically rather than assuming:

  • Where is data stored and processed, and does that meet your organization’s data residency requirements?
  • What de-identification process is used before annotators see the data, and is it applied consistently across all data types (images, video, audio)?
  • Is there a signed BAA (Business Associate Agreement) available, and does the vendor’s subcontracting model (if any) extend the same obligations downstream?
  • Does the annotation workforce operate under documented data-handling agreements, or informal terms?

5. Cost structure and hidden costs

Medical annotation vendors often structure pricing differently. Rather than comparing the headline price alone, ask vendors exactly what the quote covers. 

  • What’s the cost basis – per annotation hour, per image/item, per project?
  • How does cost scale with specialty requirements (a subspecialist’s time costs more than a generalist’s, a vendor whose pricing doesn’t reflect that isn’t actually using subspecialists)?
  • What’s excluded from the quoted price: adjudication, QA sampling, documentation, revisions?
  • What does rework typically cost if initial labels don’t meet your quality bar, and how is that measured?

The cheapest quote is rarely the cheapest project once you account for rework, delayed submissions, or a dataset that fails an internal audit six months in.

For a meaningful comparison, evaluate the total cost of getting a dataset that meets your required quality, documentation, security, and overall SLA’s, rather than the price of an individual annotation alone.

Red Flags When Choosing a Medical Data Annotation Vendor

  • Vague answers to the specialty-matching question, or a claim that one “medical annotator” pool covers every modality equally well.
  • No willingness to share agreement/quality metrics, even in aggregate or anonymized form.
  • No documented process for handling annotator disagreement.
  • Pricing that doesn’t vary by specialty or complexity, a sign that specialty matching isn’t actually happening.
  • Reluctance to discuss data provenance or documentation for regulatory use, or a generic “we’re compliant” answer with no specifics.

How we'd answer this checklist ourselves

Since we’ve emphasized the importance of avoiding vague answers, here’s what a specific, well‑structured response looks like, using our own process as an example.

  • Credentials and specialty matching: every dataset is assigned based on scan type, clinical domain, and annotator track record with similar data, radiology to radiologists, pathology to pathologists, and so on, then reviewed by a licensed specialist in that same domain before delivery.
  • Quality methodology: inter-annotator agreement scores are included with every delivered dataset by default, not on request, described internally as “ready for regulatory review”, not a retrofit, part of standard delivery.
  • Documentation: datasets ship in your required format (DICOM, NIfTI, NRRD, COCO, Pascal VOC, or your proprietary platform’s format) alongside that QA documentation.
  • Security: GDPR-compliant as an EU-based provider; HIPAA posture depends on whether your data contains PHI, which we’ll confirm against a sample before starting.
  • Cost and commitment: we run a free, no-commitment trial round on a sample of your data before any pricing conversation, so you can evaluate the actual annotators and output quality first.

At Doctors in the Loop, we hold every project to this standard, because we chose not to compete on price. We compete on the opposite axis: fair pay, safe work, and real training, because that’s what produces medical AI training data that teams can actually trust. Book a call to talk through your dataset.

Frequently asked questions

What questions should be in an RFP for medical data annotation? At minimum: annotator credential verification and specialty-matching process, inter-annotator agreement methodology, adjudication workflow for disagreements, documentation available for regulatory use (EU AI Act data governance, FDA GMLP), data security and de-identification process, and a transparent cost structure by specialty and complexity.

How do I verify an annotation vendor actually uses licensed clinicians? Ask for specialty-matched confirmation tied to your specific project, not a general claim about the vendor’s workforce. A vendor with genuine clinical annotators will be able to confirm, at minimum, the relevant specialty of the professionals assigned to your data.

What’s a red flag when evaluating a medical annotation vendor? Vagueness on measurable quality metrics, specifically inter-annotator agreement, combined with vagueness on how disagreements between annotators are resolved. Both are basic QA infrastructure; a vendor without clear answers likely doesn’t have a mature process, regardless of marketing claims.

Does annotation documentation matter if we’re not seeking FDA clearance yet? Yes, retrofitting documentation later, once you do pursue clearance or face a conformity assessment, is far more expensive than requiring it from the start. Provenance and credential documentation should be a standing requirement, not something requested only when a regulatory deadline appears.

How Doctors in the Loop Approaches Medical Annotation

At Doctors in the Loop, this checklist maps directly onto how every project actually runs. We sign an NDA before any data moves, and it stays in your environment until you decide otherwise. Specialist matching happens next, a team is assembled based on your modality, specialty, and project goals, and you see who’s assigned before any annotation starts. 

From there we run a collaborative pilot round together to validate guidelines and confirm scope before full production begins. Delivery includes full QA documentation in COCO, DICOM, Pascal VOC, or your own format, with inter-annotator agreement scores included and ready for regulatory review. Book a call to talk through your dataset

Leave a Comment

Your email address will not be published. Required fields are marked *