Skip to content

Scale AI’s New Benchmark Finds Humans Score 93% on Everyday Visual Judgments While the Best AI Model, GPT-6 Astra, Gets 54%

Scale AI and Elorian's Humanity's Sixth Sense benchmark asks everyday visual questions. People score 93.1%, the best model, GPT-6 Astra, scores 53.6%, and GPT-6 Luna, the model OpenAI uses for free ChatGPT, comes last of 25.

Law library shelf with Brooklyn Law Review volumes and visible gaps, an example scene from the Humanity's Sixth Sense benchmark

Scale AI and the visual AI startup Elorian have released a new test called Humanity’s Sixth Sense, and the headline number is blunt. People scored 93.1%. The best AI model, OpenAI’s GPT-6 Astra running at maximum reasoning effort, scored 53.6%. The median model among the 25 tested scored 30.9%.

The questions are not hard in the usual sense. Would two more books fit on this shelf? Which way is the wind blowing, judging by the flag? Why did the runner slow down? Most adults answer these at a glance, and the same models that ace math and coding exams keep getting them wrong.

The benchmark went live on October 7 (US time) with a public leaderboard, a research paper and the full dataset on Hugging Face under an MIT license. We went through the paper’s tables rather than the launch posts, and a few details matter more than the headline.

What Scale and Elorian actually tested

The benchmark has 522 open-ended tasks: 288 built on still images and 234 on video clips, adding up to 17.6 hours of footage. Each task pairs a scene with a question written by a trained annotator, plus a reference answer and a short rubric.

Scale’s blog says every task had to follow three rules. It must require inference beyond what is visible, it must need no specialist knowledge, and the answer must get unanimous agreement from human reviewers. Of 3,466 tasks written, only 522 survived three rounds of review, an acceptance rate of 15.1%.

The tasks fall into four groups: what happened before or after the moment shown, physical and spatial logic (will it fit, can someone reach it), social understanding (who is in charge, what someone knows), and abstract patterns. There is no multiple choice. Models answer in free text, and a task only counts as solved when every rubric point is met.

The scoreboard, including the models in ChatGPT

Each model got three attempts per task, and the score is the average pass rate. Here is a selection from the paper’s main results table.

Model (reasoning effort) Vendor Overall Social tasks
Human baseline (20 people) None 93.1% 89.1%
GPT-6 Astra (max) OpenAI 53.6% 46.1%
GPT-6.1 Sol (max) OpenAI 46.6% 36.4%
Claude Opus 5.5 (xhigh) Anthropic 44.6% 34.8%
Gemini 3.8 Flash (high) Google 41.6% 40.0%
Muse Spark 1.3 (max) Meta’s Muse 37.4% 25.8%
Qwen3.8 Max (max) Qwen 33.0% 21.0%
GPT-6 Sol (max) OpenAI 31.2% 22.4%
Kimi K3 (max) Moonshot 25.5% 15.5%
GPT-6 Luna (max) OpenAI 21.0% 9.4%

The bottom row is the one ChatGPT users should notice. OpenAI’s own GPT-6 launch post says ChatGPT runs on GPT-6 Sol for paid plans and GPT-6 Luna for Free and Go users. Models with those names scored 31.2% and 21.0% here, and Luna came last of all 25. On social questions it managed 9.4%. We covered that rollout in our piece on GPT-6 reaching free ChatGPT users.

Only five of the 25 models cleared 40%. Social understanding was the weakest area for 21 of them, and video was harder than still images for 23 of them, by 7.3 points on average.

A model that “saturates” FrontierMath still misjudges a bookshelf

The contrast with OpenAI’s own claims is hard to miss. OpenAI’s GPT-6 Astra announcement says the model scores 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3. This week the company also published hundreds of math papers from an unreleased model, which we looked at in our report on OpenAI’s 722 math manuscripts.

Yet on the HSS shelf photo, Elorian’s write-up says GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “no, there is too little space” on all three attempts. The correct answer is yes, using gaps that are already visible. The same three models also misread which edge of a painting a traveler entered from, whether plough chains were taut, and which way a flag was blowing.

“I watched top AI models insist a bookshelf with plenty of space had no room for another book,” Elorian’s Andrew Dai wrote on X. “They can write professional code but miss what most people see in a second.”

More thinking does not fix it

The paper’s failure analysis is the most useful part. The team labeled 8,573 failed answers and found 53% came from perception and 41% from what it calls latent inference, meaning things implied but not shown. Only 5% were logic errors.

The two biggest causes were missing the one visual cue that decides the answer (21%) and misidentifying an object, person or role (20%). The failures were also shared across vendors: on tasks that five or more models failed, a median of 80% failed for the same reason.

Extra reasoning did not rescue them. Models burned about 4,000 reasoning tokens per task on average, and Scale says GPT-6 Astra dropped 14 points on “what already happened” questions when moved from high to extra high effort. With the image or video removed entirely, Astra fell to 6.6%, which suggests the test cannot be gamed from text alone.

Agents with zoom and crop help, but only so much

The researchers also ran models inside Claude Code and Codex on a 388-task subset, where they could crop, zoom, search the web and re-sample frames.

Harness and model Single pass As an agent Median steps
Claude Code with Claude Opus 5 30.8% 51.3% 72
Claude Code with Claude Fable 5.1 41.2% 57.7% 57
Codex with GPT-6 Astra 54.4% 59.3% 14
Codex with GPT-5.6 Sol 28.4% 35.8% 18

Claude Code added 16.5 to 20.5 points, but at more than 42,000 thinking tokens per task. The best setup still topped out at 59.3%. Zooming fixed missed cues, but errors from reading a flat 2D overlap as 3D alignment barely moved, from 102 to 101.

What is solid and what needs a pinch of salt

Solid Treat with caution
Dataset and grading code are public, so anyone can rerun it Elorian is building visual reasoning models, so a test showing rivals struggle helps its pitch
Tasks passed three independent human review rounds Answers are graded by an AI judge (Claude Opus 5, per the leaderboard)
20 humans took the test under the same free-form rules Images come from the public web and may have been seen in training
In a five-model re-grade, rankings held with judges from three vendors, Scale says Video clips had no audio or transcript
Confidence intervals are published for every score Newer models like Gemini 4 Argon and Claude Haiku 5.5 are not on the board

The conflict of interest deserves a plain mention. Elorian’s own site says it is building models that “natively understand and reason through the visual medium” and lists $55 million in backing from Striker Ventures, Menlo Ventures and Altimeter, with NVIDIA participating. Scale sells training data to AI labs. Neither fact makes the numbers wrong, and the open dataset is a fair answer to it, but it is context the launch posts skip.

Google’s newest model is also missing. Gemini 4 Argon was announced on September 30 with early access limited to security partners, as we noted when Google announced it, so the Google entries here are Flash models.

What it means for Indian developers

If you are building anything that acts on camera input, this is a test worth running against your own use case. Retail shelf audits, warehouse robots, dashcam incident review, insurance claim photos and CCTV alerts all lean on exactly the “will it fit” and “what just happened” judgments where models scored worst.

The practical takeaway is to keep a human check on visual decisions that carry money or safety risk, and to test with your own photos rather than trusting a model’s general benchmark scores. The HSS dataset is small enough at 522 tasks to run on a modest budget, and the license allows commercial use.

The same goes for teams in the US, Canada and Australia shipping vision features in logistics, retail or home robotics. A model that writes great code can still miss the decisive detail in a photo.

The bottom line

Humanity’s Sixth Sense does not show that AI is useless at vision. It shows that the gap sits in a specific place: everyday intuition about space, cause and people. That is the layer robots, cars and home assistants need most, and right now even the strongest model gets it right only about half the time.

FAQ

What is the Humanity’s Sixth Sense benchmark?

It is a test of intuitive visual reasoning built by Scale AI and Elorian. It has 522 open-ended questions about images and video that most people can answer at a glance, such as whether an object will fit or why someone is reacting.

Which AI model scored highest?

OpenAI’s GPT-6 Astra at maximum reasoning effort scored 53.6%, followed by GPT-6.1 Sol at 46.6% and Claude Opus 5.5 at 44.6%. Humans scored 93.1%.

How did the models in ChatGPT do?

GPT-6 Sol, the name of the model behind ChatGPT’s paid plans, scored 31.2%. GPT-6 Luna, used for Free and Go users, scored 21.0% and ranked last of the 25 models tested.

Can I run the benchmark myself?

Yes. Scale has published the dataset on Hugging Face under an MIT license, along with its evaluation code.

Is the benchmark neutral?

It is open and reviewed, but not free of interests. Elorian is building visual reasoning models of its own, answers are scored by an AI judge, and some images may have appeared in training data. Scale says that when it re-graded five models with judges from three vendors, the rankings did not change.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *

Loading the next article…

Continue reading