The integration of AI-assisted biomedical image analysis into clinical practice demands AI-generated findings that are not only accurate but also interpretable. However, existing models generally lack the ability to simultaneously generate diagnostic findings and localize corresponding targets. This limitation makes it challenging to correlate AI-generated findings with visual evidence and interpret the results. To this end, we introduce UniBiomed, a universal foundation model for grounded biomedical image interpretation, which is capable of generating accurate diagnostic findings and segmenting the biomedical targets. UniBiomed is based on an integration of Multi-modal Large Language Model and Segment Anything Model, which can unify diverse biomedical tasks in universal training for advancing grounded interpretation. To develop UniBiomed, we curate a large-scale dataset comprising 27 million triplets of images, region annotations, and text descriptions. Extensive validation on 70 internal and 14 external datasets demonstrated the state-of-the-art performance of UniBiomed in diverse biomedical tasks.
© 2026. The Author(s).