{"slug":"benchmarking-inference","label":"Benchmarking Inference","item_count":3,"day_count":3,"source_count":3,"first_seen":"2026-09-08T16:21:57+00:00","last_updated":"2026-09-28T09:09:42+00:00","generated_at":"2026-09-29T15:25:23.247370+00:00","sources":["arxiv_agent_systems_research","aws_ml_blog","hackernews_ai"],"days":[{"date":"2026-09-08","items":[{"title":"Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6","url":"https://aws.amazon.com/blogs/machine-learning/benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6","source":"aws_ml_blog","type":"news","summary_1line":"Benchmark two 30B Mixture-of-Experts models, Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B, across G5, G6, G6e, and G7 GPU instances on Amazon SageMaker AI. Compare throughput, latency, and cost-per-token, and see ho...","sid":"8b3a4550ce0344d8","published":"2026-09-08T16:21:57+00:00","editor_note":"Instance-level baseline: throughput, latency and cost-per-token for two 30B MoE models across four GPU generations."}]},{"date":"2026-09-24","items":[{"title":"SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving","url":"https://research.nvidia.com/benchmarks/swe-serve","source":"hackernews_ai","type":"news","summary_1line":"SWE-Serve: Benchmarking Agentic Engineering for Production Inference Serving","why_it_matters":"Matches feed focus: agentic.","sid":"485a8d96b72d9981","published":"2026-09-24T23:34:34+00:00","editor_note":"Shifts the benchmark target from raw instance speed to agentic engineering workloads on production serving."}]},{"date":"2026-09-28","items":[{"title":"AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs","url":"http://arxiv.org/abs/2609.34683v1","source":"arxiv_agent_systems_research","type":"paper","summary_1line":"The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. Howeve...","why_it_matters":"Matches feed focus: agentic, evaluation.","sid":"8424df2cc3f5d4ac","published":"2026-09-28T09:09:42+00:00","editor_note":"Academic benchmark suite specifically for agentic-LLM inference performance in serving engines."}]}],"editorial":{"tldr":"On Sep 8 AWS published GPU-instance benchmarks for small LLM inference on SageMaker AI, comparing G5, G6, G6e and G7 on throughput, latency and cost per token. Sep 24 brought NVIDIA's SWE-Serve, which benchmarks agentic engineering for production inference serving.","stale":false,"whats_new":"AgentPerfBench (arXiv, Sep 28) proposes a benchmark suite for inference performance of agentic LLMs, arguing that serving engines like vLLM and SGLang are tuned against workloads that don't represent agents.","why_it_matters":"Serving-engine and hardware choices are made against benchmark workloads, so instance sizing and scheduler tuning based on chat-style traffic may not hold for agent traffic.","take_for_builders":"Before sizing GPU instances from chat-workload benchmarks, replay your own agent traces (long prefixes, many tool-call turns) through vLLM or SGLang and compare cost per token.","status":{"state":"Developing","tone":"rising","changed":"2026-09-28","detail":"Three benchmarking efforts in three weeks, moving from instance-level GPU comparisons toward agent-specific serving workloads."},"beats":[{"kicker":"BASELINE","tone":"launch","headline":"AWS benchmarks Qwen3-Coder-30B and Nemotron-3-Nano-30B across G5, G6, G6e and G7","summary":"Instance-level comparison of throughput, latency and cost per token for small MoE models.","sids":["8b3a4550ce0344d8"]},{"kicker":"SERVING","tone":"rising","headline":"NVIDIA publishes SWE-Serve for agentic engineering on production inference serving","sids":["485a8d96b72d9981"]},{"kicker":"NOW","tone":"now","headline":"AgentPerfBench targets inference performance of agentic LLMs on vLLM and SGLang-style engines","summary":"Argues serving optimizations are benchmark-driven, so representative agent workloads matter.","sids":["8424df2cc3f5d4ac"]}],"open_questions":["Do AgentPerfBench results show different engine or hardware rankings than chat-style benchmarks?","Will AWS repeat its G7 vs G5/G6 comparison with an agentic workload?"],"generated_at":"2026-09-29T05:30:00+00:00"}}