arXivarxivObjective vs. Search: Decomposing What Makes a Good Tokeniserarxivcs.CLcs.AI5 days ago▲ 43Read story→
arXivarxivA Zeroth-Order Paradigm for LLM Preference Alignmentarxivcs.CLcs.AI5 days ago▲ 43Read story→
arXivarxivPANORAMA: Panoptic Grounded Captioning via Mask Proposal Selectionarxivcs.CVcs.CL5 days ago▲ 43Read story→
arXivarxivDreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generationarxivcs.ROcs.AI5 days ago▲ 43Read story→
arXivarxivExponential Hardness of Off-Policy Evaluation under History-Dependent Loggingarxivcs.LGpublisher:arxiv5 days ago▲ 43Read story→
arXivarxivScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environmentsarxivcs.CLcs.CY5 days ago▲ 43Read story→
arXivarxivCognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environmentsarxivcs.AIcs.LG5 days ago▲ 43Read story→
arXivarxivAffora: A Design System for Agent-Friendly Interfacesarxivcs.HCcs.AI5 days ago▲ 43Read story→
arXivarxivFlag Game: A Toy Model for Mechanistic Swarm Interpretabilityarxivcs.AIcond-mat.dis-nn5 days ago▲ 43Read story→
arXivarxivPlaying log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Modelsarxivcs.CLpublisher:arxiv5 days ago▲ 43Read story→
arXivarxivHow Model Growth, Recursion, and Boundary Operators Influence Scaling Exponentsarxivcs.LGpublisher:arxiv5 days ago▲ 43Read story→
arXivarxivrMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inferencearxivcs.ROcs.AI5 days ago▲ 43Read story→
arXivarxivMonitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluationsarxivcs.CLcs.LG5 days ago▲ 43Read story→
arXivarxivEvidence-Grounded Agentic Formulation Development in an Autonomous Laboratoryarxivcs.LGpublisher:arxiv5 days ago▲ 43Read story→
arXivarxivPrepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeriaarxivcs.CYcs.AI5 days ago▲ 43Read story→
arXivarxivReporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluationarxivcs.CLcs.AI5 days ago▲ 43Read story→
arXivarxivSecuring quantum error correction against misleading advice from AI agentsarxivquant-phcs.AI5 days ago▲ 43Read story→
arXivarxivMUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Educationarxivcs.AIcs.CL5 days ago▲ 43Read story→
arXivarxivA General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddingsarxivstat.MLcs.LG5 days ago▲ 43Read story→