arXivarxivConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Modelsarxivcs.CLpublisher:arxivabout 1 month ago▲ 43Read story→
arXivarxivG-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretationarxivcs.CLcs.AIabout 1 month ago▲ 43Read story→
arXivarxiv$TCP_α$: Margin-Controlled Confidence estimation for reliable Music Information Retrievalarxiveess.AScs.LGabout 1 month ago▲ 43Read story→
arXivarxivA comparison between ceiling-mounted FMCW, IR-UWB and Wi-Fi radar for in-bedroom human activity monitoring and sleep interruption detectionarxivcs.LGpublisher:arxivabout 1 month ago▲ 43Read story→
arXivarxivAn Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Predictionarxivcs.AIcs.CLabout 1 month ago▲ 43Read story→
arXivarxivInducing Task Models from Computer-Use Tracesarxivcs.CLcs.AIabout 1 month ago▲ 43Read story→
arXivarxivAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvementarxivcs.AIcs.CLabout 1 month ago▲ 43Read story→
arXivarxivPandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimationarxivcs.AIpublisher:arxivabout 1 month ago▲ 43Read story→
arXivarxivExplainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Recordsarxivcs.LGpublisher:arxivabout 1 month ago▲ 43Read story→
arXivarxivMidTool: Mid-training Data Synthesis for Agentic Tool Usearxivcs.AIpublisher:arxivabout 1 month ago▲ 43Read story→
arXivarxivPhysical-Support Confidence Sets for Highly Coherent Dictionariesarxivcs.LGeess.SPabout 1 month ago▲ 43Read story→
arXivarxivPhantom Gains: Auditing Self-Improvement Against a Measured Nullarxivcs.AIcs.CLabout 1 month ago▲ 43Read story→
arXivarxivDynamic Structural Causal Modeling for Sleeparxivcs.LGpublisher:arxivabout 1 month ago▲ 43Read story→
arXivarxivInject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalizationarxivcs.CLcs.AIabout 1 month ago▲ 43Read story→
arXivarxivWhich Eviction Policy Should an LLM Cache Use? A Systematic Study Across Workloads, Capacities, and Encodersarxivcs.DBcs.LGabout 1 month ago▲ 43Read story→
arXivarxivBreak It Down, Pass It On: Cross-Task Skill Transfer in LLM Agentsarxivcs.AIcs.CLabout 1 month ago▲ 43Read story→
arXivarxivCatching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learningarxivcs.AIcs.DCabout 1 month ago▲ 43Read story→
arXivarxivDICS: Data-Informed Centroid Splitting for Decision Tree Classifiersarxivcs.LGstat.MLabout 1 month ago▲ 43Read story→
arXivarxivLearning When to Think: Adaptive Reasoning for Test-Time Compute Allocationarxivcs.AIpublisher:arxivabout 1 month ago▲ 43Read story→