Introducing Align Evals: Streamlining LLM Application Evaluation
Align Evals is a new feature in LangSmith that helps you calibrate your evaluators to better match human preferences. Context & related coverage →
AI news for platform & agent engineers
Ranked signal · finite reading
One shared ranking. Scan what changed, save what matters, and stop when the finish line appears.
Loading...
Align Evals is a new feature in LangSmith that helps you calibrate your evaluators to better match human preferences. Context & related coverage →
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows... Context & related coverage →
In today’s agentic era, modern cloud applications are evolving from a set of passive tools to fleets of autonomous digital workers that reason, plan, and take action across a wide range of tasks. For platform engineer... Context & related coverage →
Introducing LangSmith LLM Gateway: runtime governance for AI agents with spend limits, PII redaction, and trace continuity, built directly into LangSmith. Context & related coverage →
Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously dis... Context & related coverage →
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable... Context & related coverage →
release(checkpoint-sqlite): 3.1.1 · fix(checkpoint-postgres,checkpoint-sqlite): scope namespace matching to segment boundaries · docs: standardize package README.md structure Context & related coverage →
The writing assignments I give my students are gym tasks, not work tasks. I ask them to write policy memos not because the world needs more policy memos. I assign them because the very act of writing, which includes t... Context & related coverage →
Added a streaming tool-call parse buffer limit to the OpenAI-compatible frontend to prevent excessive memory usage during streaming tool calls · PyTorch backend: Enabled PyTorch 2 batching and added support for loadin... Context & related coverage →
DeepSeek is developing a massive AI data centre in Inner Mongolia AFR Context & related coverage →
AI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries. Context & related coverage →