Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
2024 · 2 citations
https://doi.org/10.52202/079017-2400
Claim or correct this author profile
Claim this profile or suggest corrections to the name, affiliation, bio, photo, or paper titles. Approved changes appear as verified CitedEvidence overlays.
74
Papers
157
Citations
4
h-index
3
i10-index
Zhenmei Shi is an academic researcher from University of Wisconsin–Madison. The author has contributed to research in topics: Neural Networks and Applications & Natural Language Processing Techniques & Multimodal Machine Learning Applications. The author has an h-index of 4, co-authored 63 publications.
ORCID: 0009-0007-6741-7598https://doi.org/10.52202/079017-2400
https://doi.org/10.48550/arxiv.2502.16490
https://doi.org/10.48550/arxiv.2503.14881
Click to start Chat