AtlasGPT: a language model grounded in neurosurgery with domain-specific data and document retrieval

J Neurosurg 143:560–567, 2025

AtlasGPT, a neurosurgery-specific large language model grounded in expert-verified sources and retrieval-augmented generation, outperformed GPT-4 and Gemini Advanced on a neurosurgery board exam, showed greater resistance to medical misinformation, and generated more comprehensive, relevant, and well-referenced answer explanations than standard preparation materials.

• AtlasGPT is a neurosurgery-specific large language model (LLM) built on GPT-4 with retrieval-augmented generation (RAG) from trusted neurosurgical sources.

• AtlasGPT outperformed GPT-4 and Gemini Advanced on a 149-question neurosurgery board exam (accuracy: 90.6% vs 80.5%).

• AtlasGPT showed the highest accuracy on spine and imaging-based questions, even without access to image data.

• In adversarial testing, AtlasGPT was more robust to misinformation, being fooled only 14% of the time, compared to 44% for GPT-4 and 68% for Gemini Advanced.

• Expert neurosurgeons rated AtlasGPT’s explanations as more comprehensive, relevant, and better referenced than official board prep materials.

• AtlasGPT did not produce hallucinations or harmful content in its responses.

• The study suggests domain-specific LLMs like AtlasGPT can enhance medical education, decision-making, and exam preparation in complex fields.

• Limitations include use of a single question bank and need for broader source material in future work.