Shelf

Practice

What would a team do differently this quarter?

Work an engineering team could act on now.

20 articles, newest first
ArticlePublisherSurfaced
September 2026 4 articles
New Benchmark Finds Agents Save ERP Records Correctly in as Few as 3% of Runs arXiv cs.AI
Pareto Atlas Maps Which LLM Inference Optimizations Actually Win arXiv cs.AI
LLM-Judge Satisfaction Scores Fail to Predict Task-Oriented Agent Success arXiv cs.CL
Raw Bash Access Beats Typed Tool Interfaces for Enterprise Agents arXiv cs.SE
August 2026 16 articles
License-Aware Distillation Recipe for CPU-Deployable Safety Classifiers arXiv cs.AI
IBM Granite 4.2: Dense Reasoning LLMs at 3B, 8B, and 30B Hugging Face Blog
F2Asm Learns Exact SASS Assemblers for Blackwell and Rubin GPUs arXiv cs.LG
When Fine-Tuned Classifiers Beat LLMs at Intent Detection, and When They Don't arXiv cs.CL
Meta Details Its Deployment Health Check System for Automatic Rollback arXiv cs.SE
Study Compares Copilot, Cursor, Windsurf on Full-Stack App Generation arXiv cs.SE
Rust Glancer: New Rust LSP Claims 100x Lower RAM Than rust-analyzer matklad
Nari Labs cuts Qwen3-TTS time-to-first-audio to 34ms on H100 toebee
Timing Experiments Reverse-Engineer NVIDIA GPU Memory Access Paths ibobev
Building a 128-GPU Cluster from Retired Hardware for LLaMA-70B Inference arXiv cs.LG
DAG-Structured Multi-Agent System Fixes LLM Failures on Clinical Trial Coding arXiv cs.AI
LLM Prompting Recovers Institution-Specific PHI Missed by De-identification Tools arXiv cs.CL
LiquidAI Releases Q4_0 GGUFs Trained via Quantization-Aware Distillation Hugging Face Blog
Spec-First Prompting Improves LLM Test Generation on Production Bugs arXiv cs.SE
MongoDB's 11-Year Push to Unify Conformance Tests Across a Dozen SDKs arXiv cs.SE
TEMPO Balances MoE Expert Dispatch Across Memory and Compute Regimes arXiv cs.CL

Back to the latest edition How articles are chosen