Doximity published Bedside Bench, an open source clinical AI benchmark of 500 cases across 10 sub benchmarks, including drug safety, guideline adherence, landmark trials, hallucination and false premises, health equity and diagnostic safety. The 250 case training split and the grading rubrics are public on Hugging Face, and the test split is held out. Bedside Bench is one of the domain benchmarks in Fireworks' Specialized Intelligence Index, and Doximity states that Doximity Ask ranked first against frontier models when Fireworks graded it.
Our readBuyers get a public rubric they can run against Doximity Ask and competing clinical reference tools. The ranking is stated without published scores and the benchmark was written by Doximity, so ask for the per benchmark numbers before relying on it.