Skip to content

Speaker: Katherine Xu – Language and Vision Working Group

Virtual

Title: Are Vision-Language Models Checking or Looking? Abstract: Today’s AI vision systems are trained on vast amounts of data, yet it remains unclear whether they simply retrieve memorized answers or actively reason. We conjecture that hallucinations and limited creativity in these models stem from an over-reliance on superficial "checking" rather than active "looking." Checking retrieves the…