J Neurosurg 143:560–567, 2025
AtlasGPT, a neurosurgery-specific large language model grounded in expert-verified sources and retrieval-augmented generation, outperformed GPT-4 and Gemini Advanced on a neurosurgery board exam, showed greater resistance to medical misinformation, and generated more comprehensive, relevant, and well-referenced answer explanations than standard preparation materials.
• AtlasGPT is a neurosurgery-specific large language model (LLM) built on GPT-4 with retrieval-augmented generation (RAG) from trusted neurosurgical sources.
• AtlasGPT outperformed GPT-4 and Gemini Advanced on a 149-question neurosurgery board exam (accuracy: 90.6% vs 80.5%).
• AtlasGPT showed the highest accuracy on spine and imaging-based questions, even without access to image data.
• In adversarial testing, AtlasGPT was more robust to misinformation, being fooled only 14% of the time, compared to 44% for GPT-4 and 68% for Gemini Advanced.
• Expert neurosurgeons rated AtlasGPT’s explanations as more comprehensive, relevant, and better referenced than official board prep materials.
• AtlasGPT did not produce hallucinations or harmful content in its responses.
• The study suggests domain-specific LLMs like AtlasGPT can enhance medical education, decision-making, and exam preparation in complex fields.
• Limitations include use of a single question bank and need for broader source material in future work.

You must be logged in to post a comment.