FHIR-AgentBench

AI benchmark from Paper Arxiv

FHIR-AgentBench evaluates LLM agents on realistic interoperable EHR question answering over HL7 FHIR resources. It grounds 2,931 real-world clinical questions in FHIR and compares retrieval strategies, interaction patterns, and reasoning approaches such as direct FHIR API calls, specialized tools, single-turn versus multi-turn interaction, and natural-language versus code-generation reasoning.

Publisher
Paper Arxiv
Domains
Healthcare
Data source
benchmarklist
Catalogued
Website
https://benchmarklist.com/benchmarks/fhir_agentbench/

Catalogue entry from the RL Engineering daily scrape of public sources. Data reflects the 2026-08-30 snapshot.