Shawn Pickett

Radiation Safety Director & Full-Stack Developer.

I bridge the gap between clinical operations and modern technology. Currently Director of Radiation Safety Support Services at Landauer, where I scale service lines, automate regulatory workflows, and drive profitability through technical innovation.

Shawn Pickett

Focus

Strategy • Automation • Scale

Domain

Healthcare • Safety • Regulations

Published Research

Journal of Nuclear Medicine TechnologyMay 2025

Comparison of Large Language Models’ Performance on 600 Nuclear Medicine Technology Board Examination–Style Questions

Michael A. Oumano & Shawn M. Pickett

This study investigated the application of large language models (LLMs) with and without retrieval-augmented generation (RAG) in nuclear medicine, particularly their performance across various topics relevant to the field, to evaluate their potential use as reliable tools for professional education and clinical decision-making. Methods: We evaluated the performance of LLMs, including the OpenAI GPT-4o series, Google Gemini, Cohere, Anthropic, and Meta Llama3, across 15 nuclear medicine topics. The models’ accuracy was assessed using a set of 600 sample questions, covering a range of clinical and technical domains in nuclear medicine. Overall accuracy was measured by averaging performance across these topics. Additional performance comparisons were conducted across individual models. Results: OpenAI’s models, particularly openai_nvidia_gpt-4o_final and openai_mxbai_gpt-4o_final, demonstrated the highest overall accuracy, achieving scores of 0.787 and 0.783, respectively, when RAG was implemented. Anthropic Opus and Google Gemini 1.5 Pro followed closely, with competitive overall accuracy scores with RAG. Cohere and Llama3 models showed more variability in performance, with the Llama3 ollama_llama3 model (without RAG) achieving the lowest accuracy. Discrepancies were noted in question interpretation, particularly in complex clinical guidelines and imaging-based queries. Conclusion: LLMs show promising potential in nuclear medicine, improving diagnostic accuracy, especially in areas like radiation safety and skeletal system scintigraphy. This study also demonstrates that adding a RAG workflow can increase the accuracy of an off-the-shelf model. However, challenges persist in handling nuanced guidelines and visual data, emphasizing the need for further optimization in LLMs for medical applications.

Wilderness & Environmental MedicineApril 2025

Evaluating Large Language Models on Aerospace Medicine Principles

Kyle D. Anderson, MD, PhD; Cole A. Davis, BS; Shawn M. Pickett, BS, MBA; Michael S. Pohlen, MD

Large language models (LLMs) hold immense potential as clinical decision-support tools for Earth-independent medical operations. However, the generation of incorrect information may be misleading or even harmful when applied to care in this setting. Method: This work tested two publicly available LLMs, ChatGPT-4 and Google Gemini Advanced, as well as a custom Retrieval-Augmented Generation (RAG) LLM on factual knowledge and clinical reasoning in aerospace medicine. The models were assessed for their consistency and reasoning using board-style questions. Results: ChatGPT-4, Gemini Advanced, and RAG LLMs achieved high, yet varied, scores. Nevertheless, all models exhibited gaps in factual knowledge and inconsistencies that could prove harmful. Quantitatively, the RAG LLM achieved the highest accuracy. Conclusion: There is considerable promise for LLM use in autonomous medical operations in spaceflight, but their current limitations indicate that further advancements in model training and healthcare-specific fine-tuning are required.