Introduction: Psoriasis is common, but diagnosis and early severity assessment can be delayed because of variable presentation and overlap with mimicking dermatoses. Multimodal large language models (LLMs) may assist image-based triage.
Objective: To evaluate web-based multimodal LLMs for psoriasis identification, Physician Global Assessment (PGA) scoring, and treatment recommendation quality from clinical photographs.
Methods: We retrospectively analyzed 303 standardized photographs from 160 patients (Semmelweis University, May 2022-January 2025), including 163 psoriasis lesions and 140 mimickers. Reference diagnosis and PGA were assigned by two dermatologists with third-expert adjudication, treatment outputs were rated for appropriateness. ChatGPT-5, ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.5 received identical prompts for diagnosis and, when psoriasis was predicted, PGA and treatment; sessions were periodically reset.
Results: Diagnostic accuracy was highest for ChatGPT-5 (93.1%) and ChatGPT-4o (90.1%), followed by Claude (83.6%) and Gemini (61.5%). Among correctly identified psoriasis cases, PGA accuracy was 93.3% (ChatGPT-5), 92.9% (ChatGPT-4o), 82.7% (Gemini), and 78.1% (Claude). Appropriate treatment recommendations were most frequent for ChatGPT-5 (83.7%) and ChatGPT-4o (82.1%), then Claude (75.0%) and Gemini (53.2%); test-retest agreement favored the OpenAI models.
Conclusion: Under standardized conditions, general-purpose multimodal LLMs, especially ChatGPT-5 and ChatGPT-4o, showed strong performance for psoriasis recognition and reasonable support for PGA scoring and treatment suggestions, supporting potential use by primary care physicians when dermatology specialist access is limited.
Keywords: AI; ChatGPT; Claude; Gemini; artificial intelligence; diagnosis; large language model; psoriasis.
Copyright © 2026 Boostani, Gisondi, Bellinato, Kiss, Zouboulis, Kuraitis, Goldfarb, Nádudvari, Farabi, Hodosi, Holló, Wikonkal, Lőrincz, Banvölgyi, Paragh and Kiss.