AI & Machine Learning– category –
-
AI & Machine Learning
New ‘TAM’ Benchmark Exposes Critical Gaps in GPT-5’s Long-Horizon Procedural Reasoning, Achieving Only 1% Exact Match on ICD-10-CM Clinical Coding
arXiv Unknown Overview A new benchmark, Tasks over Application Manuals (TAM), has been introduced on arXiv to accurately assess LLMs' long-horizon procedural reasoning. Evaluating GPT-5, TAM revealed alarmingly low exact match performanc... -
AI & Machine Learning
LLM Test-Time Scaling: Candidate Generation Strategy Dramatically Increases Energy Consumption by 4.86x and Latency by 6.12x on A100 GPUs for Phi-3-mini and Qwen2.5-1.5B
Mobina Kashaniyan (referencing IEEE/ACM SC26 Workshop) Unknown Overview New research reveals that for LLM test-time scaling, the candidate generation strategy, not just the candidate count, critically impacts energy and performance. Eval... -
AI & Machine Learning
LLMAR: Tuning-Free Framework Boosts Recommendation Performance in Sparse, Text-Rich Industrial Domains via LLM Inference and Self-Verification
arXiv Unknown Overview arXiv introduces LLMAR, a tuning-free recommendation framework for sparse and text-rich industrial domains, demonstrating superior accuracy, explainability, and operational cost efficiency compared to traditional t... -
AI & Machine Learning
arXiv Introduces EgoPathBench: New Benchmark Reveals Limits of Zero-Shot Egocentric Waypoint Decision-Making in Vision-Language Models
arXiv Unknown Overview A new benchmark, EgoPathBench, has been introduced on arXiv to evaluate the zero-shot egocentric waypoint decision-making capabilities of Vision-Language Models (VLMs). Comprising 31,852 training, 1,345 validation,... -
AI & Machine Learning
MedRxiv Study Reveals Mammography VLM Accuracy Overstates Evidence Grounding and Abstention Reliability
medRxiv USA Overview A study published on medRxiv indicates that accuracy metrics in mammography Vision-Language Models (VLMs) may overestimate their underlying evidence grounding and abstention reliability. The research introduces an 'e... -
AI & Machine Learning
Harrison Zhang et al. Unveil ‘Virtual Biotech’: Multi-Agent AI System to Transform Drug Discovery Decision-Making
EurekAlert! USA Overview Harrison Zhang and colleagues have introduced 'Virtual Biotech,' a novel multi-agent AI system designed to enhance drug discovery decision-making. This platform coordinates specialized AI 'scientist' agents under... -
AI & Machine Learning
Samsung Research Unveils AnySimLite: Sub-700KB, Sub-30ms On-Device AI Matching 7B-Parameter LLMs for Speech Classification
Samsung Research South Korea Overview Samsung Research has introduced AnySimLite, a lightweight few-shot similarity encoder capable of delivering state-of-the-art performance for multiple on-device speech-adjacent classification tasks. O... -
AI & Machine Learning
Φ-Bench: New Benchmark with 85 Real-World Tasks Evaluates LLMs’ Ability to Engineer and Optimize Their Own Infrastructure
arXiv USA Overview Researchers have introduced Φ-Bench (Frontier AI Infrastructure Benchmark), a novel evaluation suite comprising 85 real-world tasks designed to assess Large Language Models' (LLMs) capacity to develop and optimize thei... -
AI & Machine Learning
NovGauge: A Human-Anchored Fine-Grained Benchmark Diagnoses LLMs’ Capability in Scientific Paper Novelty Assessment, Revealing Hallucinations and Lack of Justification
arXiv USA Overview Researchers have proposed NovGauge, a fine-grained, human-anchored benchmark designed to diagnose Large Language Models' (LLMs) capabilities in assessing scientific paper novelty. Comprising 619 paper pairs and 50 mult... -
AI & Machine Learning
Stanford University Develops ‘Virtual Biotech’ with 37,000 AI Agents for Drug Discovery, Published in Science
Science (via Chosunilbo DB) USA Overview A Stanford University research team has developed 'Virtual Biotech,' a system where up to 37,000 AI agents collaborate on drug discovery, with findings published in the journal Science. This syste...