Researchers at Alibaba’s Tongyi Lab have developed an innovative framework named AgentEvolver that empowers AI agents to autonomously generate their own training data. By leveraging the reasoning capabilities of large language models (LLMs), the system explores its application environments to create synthetic, task-specific data, addressing the costly and labor-intensive nature of conventional dataset collection.
Challenges in Training AI Agents with Reinforcement Learning
Reinforcement learning (RL) is a prominent method for training AI agents to interact with environments and learn from feedback. However, applying RL to develop capable LLM-based agents presents major hurdles. Collecting relevant training datasets often demands extensive manual effort, especially when targeting novel or proprietary software lacking pre-existing data. Additionally, RL typically requires running a large volume of trial-and-error interactions, which is computationally expensive and inefficient, limiting its feasibility for bespoke enterprise applications.
AgentEvolver’s Self-Evolving Agent System
AgentEvolver introduces a paradigm shift by granting AI agents autonomy over their learning process. Described by the developers as a “self-evolving agent system,” it harnesses LLM reasoning to form a continuous self-training loop. This enables agents to interact directly with their target environments, generating tasks and refining performance without relying on predefined datasets or reward functions.
According to the research team, this approach facilitates “autonomous and efficient capability evolution through environmental interaction.” The framework’s design integrates three core mechanisms:
- Self-questioning: The agent actively explores its environment to identify functional boundaries and potential tasks, similar to how a new user experiments with software. This exploration allows the agent to produce a diverse set of training tasks tailored to user preferences, significantly reducing dependency on handcrafted datasets. Yunpeng Zhai, co-author and Alibaba researcher, explains that this mechanism transforms the agent from a “data consumer into a data producer,” cutting deployment time and costs in proprietary settings.
- Self-navigating: To enhance exploration efficiency, the agent reuses insights from previous successful and unsuccessful attempts. It learns to avoid repeating errors, such as trying to use nonexistent API functions, thus streamlining future interactions.
- Self-attributing: This mechanism provides detailed feedback by evaluating the contribution of each step within multi-step tasks. Unlike traditional RL’s sparse success/failure signals, it assesses the clarity and correctness of individual actions, which is vital for regulated industries requiring transparency and auditability.
The researchers emphasize that shifting from human-engineered training pipelines to LLM-guided self-improvement paves the way for scalable, cost-effective, and continuously advancing intelligent systems.
Technical Implementation and Scalability
AgentEvolver incorporates a Context Manager that manages the agent’s memory and interaction history, crucial for handling the complexity of real-world enterprise applications involving thousands of APIs. While acknowledging the computational challenges inherent in large action spaces, the team designed AgentEvolver’s architecture to support scalable tool reasoning adaptable to enterprise environments.
Performance Evaluation and Enterprise Implications
The framework’s effectiveness was validated using two benchmarks—AppWorld and BFCL v3—which require agents to perform complex, multi-step tasks using external tools. Using Alibaba’s Qwen2.5 family models (7B and 14B parameters), AgentEvolver was compared against a baseline trained with GRPO, a standard RL technique.
Results revealed substantial improvements: the 7B model’s average score rose by 29.4%, and the 14B model’s by 27.8%. The integration of the three mechanisms notably enhanced reasoning and task execution capabilities across both benchmarks, with self-questioning contributing the most by autonomously generating diverse training tasks and mitigating data scarcity.
Moreover, AgentEvolver demonstrated the ability to synthesize large volumes of high-quality training data efficiently, maintaining strong training performance even with limited data volume.
For enterprises, this approach offers a streamlined path to develop AI agents tailored to specific applications and workflows with minimal manual data labeling. By setting high-level objectives and allowing agents to self-generate training experiences, organizations can deploy custom AI assistants more affordably and effectively.
Future Prospects
The researchers view AgentEvolver as both a research platform and a reusable foundation for building adaptable, tool-augmented agents. Yunpeng Zhai envisions the framework as a pivotal step toward the ultimate goal: a “singular model” capable of mastering any software environment rapidly without retraining—a long-sought ambition in agentic AI.
While realizing this vision will require further advances in model reasoning and infrastructure, self-evolving systems like AgentEvolver are laying the groundwork for the next generation of intelligent, autonomous AI agents.
Fonte: ver artigo original

OpenAI Expands Executive Coaching Capabilities with Convogo Team Acquisition
xFusion Unveils Scalable Enterprise AI Hardware from Edge to Liquid-Cooled Data Centers
Meta Expands Solar Power Capacity to Fuel New AI Data Center in South Carolina
Nvidia H200 Chip Deal with China Stalls Despite Trump-Xi Summit Approval