Growing India News, world news, nation news, our news, people's news, grow news, entertainment, fashion, movies, tech, automobile and many more..
Saturday, July 12, 2025
Show HN: RULER – Easily apply RL to any agent https://ift.tt/Ch0iNza
Show HN: RULER – Easily apply RL to any agent Hey HN, Kyle here, one of the co-founders of OpenPipe. Reinforcement learning is one of the best techniques for making agents more reliable, and has been widely adopted by frontier labs. However, adoption in the outside community has been slow because it's so hard to implement. One of the biggest challenges when adapting RL to a new task is the need for a task-specific "reward function" (way of measuring success). This is often difficult to define, and requires either high-quality labeled data and/or significant domain expertise to generate. RULER is a drop-in reward function that works across different tasks without any of that complexity. It works by showing N trajectories to an LLM judge and asking it to rank them relative to each other. This sidesteps the calibration issues that plague most LLM-as-judge approaches. Combined with GRPO (which only cares about relative scores within groups), it just works (surprisingly well!). We have a full writeup on the blog, including results on 4 production tasks. On all 4 tasks, small Qwen 2.5 models trained with RULER+GRPO beat the best prompted frontier model, despite being significantly smaller and cheaper to run. Surprisingly, they even beat models trained with hand-crafted reward functions on 3/4 tasks! https://ift.tt/h2KpZOG Repo: https://ift.tt/zP3sB8K https://ift.tt/h2KpZOG July 11, 2025 at 11:17PM
Subscribe to:
Post Comments (Atom)
Show HN: Claude‑CMD – A CLI for managing Claude Code commands and workflows https://ift.tt/j98uFwp
Show HN: Claude‑CMD – A CLI for managing Claude Code commands and workflows I built Claude‑CMD, an open-source command-line interface for wo...
-
Show HN: An AI logo generator that can also generate SVG logos Hey everyone, I've spent the past 2 weeks building an AI logo generator, ...
-
Show HN: Snap Scope – Visualize Lens Focal Length Distribution from EXIF Data https://ift.tt/yrqHZtDShow HN: Snap Scope – Visualize Lens Focal Length Distribution from EXIF Data Hey HN, I built this tool because I wanted to understand which...
-
Breaking #FoxNews Alert : Biden says 'I should be the one who nominates Justice Ginsburg's successor' during campaign stop in ...
No comments:
Post a Comment