Disclosure: viggoVet operates VetEval. This post describes our side of that relationship and the rules we hold ourselves to. It does not cite or interpret any VetEval score, ranking or result. The full arrangement is set out on our VetEval and viggoVet page.
A few years into building clinical AI for veterinarians, we kept running into a question that came before all the others. We had spent a long time walking through clinics asking what an AI should do in each room: in the surgery, the dental suite, the consult room, on the road. Underneath every answer sat a harder one. Is any of this reliable enough to be in the room at all? And how would anyone outside the company that built it know?
We did not have a good answer. Nobody did. So we built one, and then we made sure we could never benefit from it directly.
The profession asked for this before we did
In March 2025 the American College of Veterinary Radiology and the European College of Veterinary Diagnostic Imaging published a joint position statement on artificial intelligence in JAVMA (vol. 263, no. 6). On the state of the market in their own specialty, it was blunt: "There is currently no commercially available product for diagnostic imaging that meets these standards" (ACVR).
The colleges went on to "urge the need for unbiased, third-party evaluations of AI tools to establish trust and ensure that these technologies meet the highest standards of clinical effectiveness." They also held that AI should always be used with a qualified veterinary professional in the loop.
That statement named the gap precisely. Veterinary AI was arriving in clinics faster than the profession could check it, and most of the claims being made about it came from the vendors making the products. Every company, including ours, had its own internal tests. None of those tests were comparable to each other, and none of them were designed to be.
What VetEval is
VetEval is a public benchmark for veterinary AI. It measures how AI models perform on real veterinary medicine using expert-authored examination items across species and specialties, reports scores with confidence intervals, and applies a safety gate that caps any model producing advice a panel of licensed veterinarians confirms as harmful. The methodology, scoring and leaderboard are published at veteval.ai.
This post leaves the leaderboard alone. We do not quote it and will not start. It covers the decision that made the benchmark worth building in the first place.
The problem with a referee who also plays
A benchmark loses its use once its operator competes on it. If viggoVet ran a benchmark and appeared on it, every result would carry an asterisk, and the asterisk would be deserved. Any good score of ours would look arranged. Any low score for a competitor would look motivated. The profession would be right to ignore the whole thing.
So the only version worth building was one where we pay a real cost. We built it, and then built a wall between it and ourselves.
So the only version worth building was one where we pay a real cost. We built it, and then built a wall between it and ourselves.
The wall, in five commitments
These are the rules VetEval runs under. Each one is written to remove a specific way the arrangement could be abused.
| Commitment | What it prevents |
|---|---|
| viggoVet models are never ranked on the VetEval public leaderboard. Not labeled, not listed, never. | The operator scoring itself |
| The item bank, gold answers and transcripts stay with the benchmark team, and are never shared with viggoVet product teams or used to train or tune anything viggoVet builds. | Teaching to the test |
| No vendor funds VetEval's operations or influences any score. Evaluation fees are uniform and published, and publication is free and decided blind to the score. | Paying for placement |
| Being a viggoVet customer, including of our APIs, has no bearing on any evaluation. Endpoints that incorporate viggoVet services are not eligible for the public leaderboard at all. | Rewarding our own customers |
| Scoring parameters and safety severity tiers are governed by licensed veterinarians and versioned with the dataset. | Moving the goalposts quietly |
The fourth commitment is the one people ask about most, because it looks commercially odd. A company that sells AI infrastructure has made its own customers ineligible for the most visible veterinary AI leaderboard. We did it because the alternative, customers who might score better for having bought from us, is exactly the conflict the whole structure exists to remove. It is simpler to make the question impossible than to answer it every time.
Each of these can be checked. The separation architecture, run conditions and dataset governance are documented at veteval.ai/methodology.
What it costs us
It costs us the most obvious piece of marketing a veterinary AI company could have.
You will not see viggoVet quote a VetEval score, embed the leaderboard, or put a benchmark badge next to a product. That absence is deliberate, and it is the policy working. It also means we give up the simple claim every AI company would like to make, a number next to our name on a public board.
We think that is the right price. Our own products are validated separately and published as viggoVet research as that work completes, with methods shown, never as leaderboard entries. A leaderboard measures what a model knows. It cannot tell you how a clinical tool behaves inside a real workflow, with a real patient and a real team, which is where most of our work actually lives.
Why a practice owner should care
If you run a practice, you are going to be asked to buy AI, probably several times this year. Most of what you will be told about it will come from the people selling it. That includes us.
The ACVR and ECVDI statement is worth reading in full for the principles alone: transparency about how a tool was built and validated, a veterinarian in the loop, and evaluation by someone other than the vendor. VetEval exists so the last of those has somewhere to happen for general veterinary AI. Its methodology is public, and the models it evaluates are named, so you can look up the model behind a product you are considering and read how it was tested. How you weigh what you find there is your call.
The short version
The profession's own specialty colleges asked for unbiased evaluation of veterinary AI. No such evaluation existed for general veterinary AI, so we built one. We then removed ourselves from it, made our own customers ineligible, and gave up quoting it, because a benchmark is only as useful as the distance between it and the people it measures.
Read more about how viggoVet and VetEval are separated, or go straight to the methodology.
References
- ACVR, (2025). Artificial Intelligence in Veterinary Diagnostic Imaging and Radiation Oncology JAVMALink
- Position statement full text:Link
- VetEval methodology:Link
- Internal: `company/VetEval/viggovet-veteval-content-package.md` (the five commitments, verbatim source)
- Disclosure line at the top and link to /veteval: present.
- No VetEval score, ranking, gate result, model performance, dataset size or provenance cited or implied.
- VetEval is not described as. independent
- No exam-style score claimed for any viggoVet product.
- Business Blog lane only: viggoVet's choices, never VetEval's findings.
- Hero: the VetEval mark on a plain background, or no image. Do not place the VetEval logo near any viggoVet product imagery.
