Benchmark– tag –
-
New Technology
University 365 Predicts 2026 Marks “The Reasoning Model Era” with GPT-6 Astra, Claude Fable 5.1 Leading Adoption
University 365 Global Overview A University 365 report states that 2026 has entered "The Reasoning Model Era," marked by the widespread adoption of reasoning models from major AI labs, including OpenAI's GPT-6 Astra, Anthropic's Claude F... -
New Technology
LLM Stats Releases September 2026 Best Reasoning AI Model Rankings, Revealing Benchmark Data for 363 Models
LLM Stats Global Overview LLM Stats unveiled its latest rankings for AI models specializing in reasoning tasks on September 17, 2026, evaluating 363 models across 444 benchmarks. This ranking provides detailed comparisons of each model's... -
New Technology
Open-Weight LLM GLM 5.3-flash Achieves High Scores on Cybersecurity Exploit Benchmarks, Industry Warned of One-Year Window to Fix Security
daily.dev (Tech Lead Digest) Unknown Overview The open-weight LLM, GLM 5.3-flash, recorded exceptionally high scores on cybersecurity exploit benchmarks like CyberGym and ExploitBench after its safety refusal capabilities were 'removed.'... -
New Technology
New ‘TAM’ Benchmark Exposes Critical Gaps in GPT-5’s Long-Horizon Procedural Reasoning, Achieving Only 1% Exact Match on ICD-10-CM Clinical Coding
arXiv Unknown Overview A new benchmark, Tasks over Application Manuals (TAM), has been introduced on arXiv to accurately assess LLMs' long-horizon procedural reasoning. Evaluating GPT-5, TAM revealed alarmingly low exact match performanc... -
New Technology
arXiv Introduces EgoPathBench: New Benchmark Reveals Limits of Zero-Shot Egocentric Waypoint Decision-Making in Vision-Language Models
arXiv Unknown Overview A new benchmark, EgoPathBench, has been introduced on arXiv to evaluate the zero-shot egocentric waypoint decision-making capabilities of Vision-Language Models (VLMs). Comprising 31,852 training, 1,345 validation,... -
New Technology
MedRxiv Study Reveals Mammography VLM Accuracy Overstates Evidence Grounding and Abstention Reliability
medRxiv USA Overview A study published on medRxiv indicates that accuracy metrics in mammography Vision-Language Models (VLMs) may overestimate their underlying evidence grounding and abstention reliability. The research introduces an 'e... -
New Technology
Φ-Bench: New Benchmark with 85 Real-World Tasks Evaluates LLMs’ Ability to Engineer and Optimize Their Own Infrastructure
arXiv USA Overview Researchers have introduced Φ-Bench (Frontier AI Infrastructure Benchmark), a novel evaluation suite comprising 85 real-world tasks designed to assess Large Language Models' (LLMs) capacity to develop and optimize thei... -
New Technology
NovGauge: A Human-Anchored Fine-Grained Benchmark Diagnoses LLMs’ Capability in Scientific Paper Novelty Assessment, Revealing Hallucinations and Lack of Justification
arXiv USA Overview Researchers have proposed NovGauge, a fine-grained, human-anchored benchmark designed to diagnose Large Language Models' (LLMs) capabilities in assessing scientific paper novelty. Comprising 619 paper pairs and 50 mult... -
New Technology
Open-Source LLMs GLM-5.2, Llama 4 Maverick, and Kimi K3 Lead Benchmarks in September 2026
Thunder Compute Global Overview In the September 2026 open-source LLM rankings, GLM-5.2, Llama 4 Maverick, and Kimi K3 demonstrated leading performance across various domains. The 744B MoE GLM-5.2 (40B parameters) excelled in reasoning b... -
New Technology
BIOSTAR Unveils NVIDIA Jetson Thor T5000/T4000-Powered Edge AI Systems at LEAP 2026, Delivering Up to 2,070 FP4 TFLOPS for Industrial AI Performance
GLOBE NEWSWIRE (via BIOSTAR) Taiwan Overview At LEAP 2026, BIOSTAR introduced a comprehensive lineup of Industrial PC (IPC) and edge AI computing solutions built on NVIDIA Jetson platforms. Notably, the MS-NAT5000 and MS-NAT4000 series, ...