MikeTrendsTrends right now

Mmastodon TechnologyAI first seen 13 h ago, last 10 h ago, peak #8

AI judge scores perfect when sums are done in code

Original: Jev got 14 maths and date questions wrong. So I did the sums in code first and let Jev judge the result: zero errors. Th

An experiment tested an AI called Jev on maths and date questions, and it got 14 of them wrong. The tester then wrote code to do the calculations and asked Jev only to judge the results, which produced zero errors. The takeaway being shared: the model was fine as a judge, the arithmetic was the weak point.

Why now: It speaks to ongoing debate about whether AI models should compute answers themselves or verify results calculated by external tools.

Jev

Open on mastodon →

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/1036440