{"slug":"reinforcement-learning","label":"Reinforcement Learning","item_count":3,"day_count":3,"source_count":2,"first_seen":"2026-07-29T06:18:16+00:00","last_updated":"2026-08-17T13:24:41+00:00","generated_at":"2026-08-19T00:13:18.629141+00:00","sources":["arxiv_cs_cl","search_cn_open_weight_labs"],"days":[{"date":"2026-07-29","items":[{"title":"Moonshot AI, kvcache-ai Open Source AgentENV To Scale Agentic Reinforcement Learning - Open Source For You","url":"https://news.google.com/rss/articles/CBMiwAFBVV95cUxQdHpXay1PcmJQUVlxRURqOGU2QlF6U3J4LW1ZVjJsNHlSWDMtVm9WRWpmZl81TnV3ZVdXZGlDOThQcUF1SUdPOXZOOEE1SjhRVmh2RkJMcldnNzlVSGVWcXhUUzAxLXN2bUtSUFNXWmxiSmd6QkI2bjNvbTlscl81cUxzdUtJeVZMUlVZeGFOMDk4U3MzTlVVaDl0Y2ZpZkszNjBMRXFfekZkbWJaZDVLRUF1RkJjYVlxQ3duOEZOUk8?oc=5","source":"search_cn_open_weight_labs","type":"news","summary_1line":"Moonshot AI, kvcache-ai Open Source AgentENV To Scale Agentic Reinforcement Learning Open Source For You","why_it_matters":"Matches feed focus: agentic.","sid":"18944a6428c3ac1a","published":"2026-07-29T06:18:16+00:00","editor_note":"Moonshot AI and kvcache-ai open-source AgentENV, an agentic-RL scaling framework — the first shipped release in this batch."}]},{"date":"2026-08-07","items":[{"title":"ResidencyRL: Reinforcement Learning in Simulated Clinical Environments","url":"http://arxiv.org/abs/2608.07418v1","source":"arxiv_cs_cl","type":"paper","summary_1line":"In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater auton...","why_it_matters":"Matches feed focus: agent, evaluation.","sid":"691438acb4bb6cca","published":"2026-08-07T17:04:41+00:00","editor_note":"New arXiv paper modeling clinical-residency training as an RL problem — physicians building expertise through staged, feedback-rich patient encounters."}]},{"date":"2026-08-17","items":[{"title":"Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning","url":"http://arxiv.org/abs/2608.16554v1","source":"arxiv_cs_cl","type":"paper","summary_1line":"Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not alwa...","why_it_matters":"Matches feed focus: evaluation.","sid":"e9e06df024881a81","published":"2026-08-17T13:24:41+00:00","editor_note":"New arXiv paper training reasoning models to ask, condition, or abstain instead of guessing when a query is missing a necessary premise, closing a known gap in answer-only RL."}]}],"editorial":{"tldr":"Reinforcement Learning is a loose weekly grouping of unrelated RL items, not a developing story. Recent entries range from Moonshot AI's open-sourced AgentENV framework for scaling agentic RL training to a paper modeling clinical-residency training as an RL problem.","stale":false,"whats_new":"A new arXiv paper trains reasoning models to ask, condition on an assumption, or abstain when a query is missing a premise needed for a unique answer — a gap in standard answer-only RL training.","why_it_matters":"Answer-only RL, the default for most reasoning-model fine-tuning, never teaches a model to flag an underspecified request; that's a real failure mode for agents built to act on ambiguous user or tool input.","take_for_builders":"RL-fine-tuning a reasoning or tool-calling model? Check whether your reward function penalizes abstaining or asking on underspecified input — answer-only RL trains models to guess rather than flag a missing premise.","beats":[{"kicker":"RELEASE","tone":"rising","headline":"Moonshot AI and kvcache-ai open-source AgentENV to scale agentic RL","summary":"New open-source framework targets scaling agentic reinforcement learning training — a shipped tool, not an academic result like the papers alongside it.","sids":["18944a6428c3ac1a"]},{"kicker":"RESEARCH","tone":"neutral","headline":"ResidencyRL simulates clinical-residency training as an RL problem","summary":"New arXiv work models how physicians convert academic knowledge into clinical judgment through simulated patient encounters with staged feedback, mirroring real residency training.","sids":["691438acb4bb6cca"]},{"kicker":"RESEARCH","tone":"neutral","headline":"New paper trains RL reasoning models to ask, condition, or abstain on missing-premise queries","summary":"Targets a known failure mode of answer-only RL: forcing a unique answer even when the query omits a premise needed to produce one.","sids":["e9e06df024881a81"]}],"open_questions":["Does AgentENV's scaling approach hold up against existing agentic-RL training frameworks, or is it mainly a Moonshot-internal tool going public?","Does ResidencyRL's simulated training transfer to real clinical decision-making, or only to the benchmark's synthetic cases?","Does the ask/condition/abstain method generalize beyond its evaluated missing-premise benchmarks to real-world ambiguous tool-calling prompts?"],"generated_at":"2026-08-19T00:20:00+00:00"}}