redditHow many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]redditmachine-learningMachineLearning2 days ago▲ 10Read story→
TMtechmemeOpenAI calls GPT-6 Astra the "world's best computer use model"; in tests, it booked DMV appointments and searched job listings faster than the average person (Maxwell Zeff/Wired)industry2 days ago▲ 24Read story→
arXivarxivSWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agentsarxivcs.SEcs.AI2 days ago▲ 43Read story→
arXivarxivFrom Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Researcharxivcs.AIpublisher:arxiv2 days ago▲ 43Read story→
arXivarxivEfficient Test-Time Adaptation through Human-AI Interactionarxivcs.AIpublisher:arxiv2 days ago▲ 43Read story→
YhackernewsLaunch HN: Mireye (YC S26) – Infrastructure for Physical World AI Agentshackernews2 days ago▲ 5Read story→
TMtechmemeSources: neocloud Fluidstack closed a $1.5B round led by Jane Street at an $18B+ valuation; in July, the company announced a $750M round at a $7.5B valuation (Iain Martin/Forbes)industry3 days ago▲ 24Read story→
arXivarxivPost-Training Language Models for Gold-Medal Performance in Coding Competitionsarxivcs.LGcs.AI3 days ago▲ 43Read story→
rssResearchers fear safety disaster ahead of OpenAI’s Astra releasethe-vergeindustryAI3 days ago▲ 44Read story→
arXivarxivOnline Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Couplingarxivcs.LGpublisher:arxiv4 days ago▲ 43Read story→
arXivarxivOrthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognitionarxivcs.CVcs.HC4 days ago▲ 43Read story→