New Technology– category –
-
New Technology
MedRxiv Study Reveals Mammography VLM Accuracy Overstates Evidence Grounding and Abstention Reliability
medRxiv USA Overview A study published on medRxiv indicates that accuracy metrics in mammography Vision-Language Models (VLMs) may overestimate their underlying evidence grounding and abstention reliability. The research introduces an 'e... -
New Technology
Harrison Zhang et al. Unveil ‘Virtual Biotech’: Multi-Agent AI System to Transform Drug Discovery Decision-Making
EurekAlert! USA Overview Harrison Zhang and colleagues have introduced 'Virtual Biotech,' a novel multi-agent AI system designed to enhance drug discovery decision-making. This platform coordinates specialized AI 'scientist' agents under... -
New Technology
Samsung Research Unveils AnySimLite: Sub-700KB, Sub-30ms On-Device AI Matching 7B-Parameter LLMs for Speech Classification
Samsung Research South Korea Overview Samsung Research has introduced AnySimLite, a lightweight few-shot similarity encoder capable of delivering state-of-the-art performance for multiple on-device speech-adjacent classification tasks. O... -
New Technology
Φ-Bench: New Benchmark with 85 Real-World Tasks Evaluates LLMs’ Ability to Engineer and Optimize Their Own Infrastructure
arXiv USA Overview Researchers have introduced Φ-Bench (Frontier AI Infrastructure Benchmark), a novel evaluation suite comprising 85 real-world tasks designed to assess Large Language Models' (LLMs) capacity to develop and optimize thei... -
New Technology
NovGauge: A Human-Anchored Fine-Grained Benchmark Diagnoses LLMs’ Capability in Scientific Paper Novelty Assessment, Revealing Hallucinations and Lack of Justification
arXiv USA Overview Researchers have proposed NovGauge, a fine-grained, human-anchored benchmark designed to diagnose Large Language Models' (LLMs) capabilities in assessing scientific paper novelty. Comprising 619 paper pairs and 50 mult... -
New Technology
Stanford University Develops ‘Virtual Biotech’ with 37,000 AI Agents for Drug Discovery, Published in Science
Science (via Chosunilbo DB) USA Overview A Stanford University research team has developed 'Virtual Biotech,' a system where up to 37,000 AI agents collaborate on drug discovery, with findings published in the journal Science. This syste... -
New Technology
US Nonprofit METR Evaluates Frontier AI Systems from Anthropic, Google, Meta, OpenAI, Focusing on Catastrophic Risks
Business Standard USA Overview METR (Model Evaluation and Threat Research), a US-based nonprofit, conducted safety evaluations of frontier AI models from leading developers including Anthropic, Google, Meta, and OpenAI. Dedicated to deve...