|
openreview.net
|
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
|
|
openreview.net
|
Contextual Drag: How Errors in the Context Affect LLM Reasoning
|
|
openreview.net
|
Learning to Continually Learn via Meta-learning Agentic Memory Designs
|
|
openreview.net
|
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
|
|
openreview.net
|
Language Self-Play For Data-Free Training
|
|
openreview.net
|
SimpleMem: Efficient Lifelong Memory for LLM Agents
|
|
openreview.net
|
Towards Execution-Grounded Automated AI Research
|
|
openreview.net
|
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
|
|
openreview.net
|
Can Language Models Discover Scaling Laws?
|
|
openreview.net
|
Self-Improving World Models via Asymmetric Forward-Inverse Consistency
|
|
openreview.net
|
Tiny Autoregressive Recursive Models
|
|
openreview.net
|
From Growing to Looping: A Unified View of Iterative Computation in LLMs
|
|
openreview.net
|
Lang-PINN: From Language to Physics-Informed Neural Networks via a Multi-Agent Framework
|
|
openreview.net
|
Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
|
|
openreview.net
|
ACE: Self-Evolving LLM Coding Framework Adversarial Unit Test Generation and Preference Optimization
|
|
openreview.net
|
Can Current Language Models Close the Discovery to Application Loop?
|
|
openreview.net
|
CausalEvolve: Towards Open-Ended Discovery with Causal Scratchpad
|
|
openreview.net
|
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
|
|
openreview.net
|
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
|
|
openreview.net
|
Test-Time Self-Distillation
|
|
openreview.net
|
Self-Evolving Rubrics: Interpretable Instance-Level Criteria for Scalable RL
|
|
openreview.net
|
Anchored Self-Play for Code Repair
|
|
openreview.net
|
Interestingness as an Inductive Heuristic for Future Compression Progress
|
|
openreview.net
|
GASP: Guided Asymmetric Self-Play For Coding LLMs
|
|
openreview.net
|
Adaptive Meta-Curriculum for Test-Time Self-Improvement
|
|
openreview.net
|
Self-Improving Clinical Reasoning via Textual Gradients
|
|
openreview.net
|
Federated Agent Reinforcement Learning
|
|
openreview.net
|
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
|
|
openreview.net
|
OMEGA: Optimizing Machine learning by Evaluating Generated Algorithms
|
|
openreview.net
|
Intelligent Robot Manipulation Requires Self-Directed Learning
|
|
openreview.net
|
Correct Reasoning Paths Visit Shared Decision Pivots
|
|
openreview.net
|
RFTF: Reinforcement Fine-tuning for Vision-language-action Models with Temporal Feedback
|
|
openreview.net
|
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
|
|
openreview.net
|
TangramSR: A Benchmark for Recursive Self-Improvement In Continuous Geometric Reasoning
|
|
openreview.net
|
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
|
|
openreview.net
|
A Framework for Prompt Optimization and Translation Across Foundation Models
|
|
openreview.net
|
Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
|
|
openreview.net
|
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
|
|
openreview.net
|
Contrastive Self-Refinement for Low-Cost Adaptation in Real-World Text-to-SQL
|
|
openreview.net
|
Simple Baselines are Competitive with Code Evolution
|
|
openreview.net
|
Rethinking Machine Unlearning: Models Designed to Forget via Key Deletion
|
|
openreview.net
|
Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
|
|
openreview.net
|
Unlocking Intrinsic Self-Reflection for LLM Preference Policy Optimization
|
|
openreview.net
|
TextBO: Bayesian Optimization in Language Space for Eval-Efficient Self-Improving AI
|
|
openreview.net
|
POLARIS: A GODEL AGENT FRAMEWORK FOR SMALL LANGUAGE MODELS THROUGH EXPERIENCE ABSTRACTED POLICY REPAIR
|
|
openreview.net
|
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
|
|
openreview.net
|
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
|
|
openreview.net
|
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
|
|
openreview.net
|
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
|
|
openreview.net
|
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
|
|
openreview.net
|
Theory-Driven Modeling and LLM-Guided Evolution for Power System Scheduling
|
|
openreview.net
|
Differentiable Evolutionary Reinforcement Learning
|
|
openreview.net
|
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
|
|
openreview.net
|
Unrolled Policy Iteration for Tiny Recursive Models
|
|
openreview.net
|
Reasoning Cache: Learning to Extrapolate to Long Lengths via Short-Length RL
|
|
openreview.net
|
ESDAE: Evaluating Synthetic Data for Agent Evaluation
|
|
openreview.net
|
Actor-Curator: Scalable Policy-driven Curriculum Learning for RL Post-Training
|
|
openreview.net
|
Reward Hacking in Self-Improving Code Agents
|
|
openreview.net
|
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
|
|
openreview.net
|
Learning What to Learn: Curriculum Curation for Test-Time Agent Learning
|
|
openreview.net
|
Beyond Solving: A Closer Look at LLMs as Solution Verifiers
|
|
openreview.net
|
Aligned but Stereotypical? Understanding and Mitigating Social Bias in LLM-Driven Text-to-Image Models
|
|
openreview.net
|
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
|
|
openreview.net
|
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
|
|
openreview.net
|
AlphaApollo: A System for Deep Agentic Reasoning
|
|
openreview.net
|
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
|
|
openreview.net
|
Leveraging Suboptimal and Noisy Trajectories for Goal-Conditional Offline RL
|
|
openreview.net
|
AUTOHARNESS: IMPROVING LLM AGENTS BY AUTOMATICALLY SYNTHESIZING A CODE HARNESS
|
|
openreview.net
|
One-Step Video Depth Estimation via Self-Distillation
|
|
openreview.net
|
Discover the distinguishing and effective reasoning patterns among LLMs via an LLM
|
|
openreview.net
|
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
|
|
openreview.net
|
Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback
|
|
openreview.net
|
Learning to Evolve: Scaling Open-Ended Discovery with Relative-Progress RL
|
|
openreview.net
|
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
|
|
openreview.net
|
Language-Guided Expertise Evolution for Protein Optimization
|
|
openreview.net
|
A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula
|
|
openreview.net
|
Adaptive Decoding via Test-Time Policy Learning for Self-Improving Generation
|
|
openreview.net
|
Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic Data
|
|
openreview.net
|
Vision-Guided Iterative Refinement for Frontend Code Generation
|
|
openreview.net
|
MAPPA: Scaling Multiagent Systems with Process Rewards
|
|
openreview.net
|
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
|
|
openreview.net
|
SAGE: Self-play Adversarial Games Enhance Large Language Model Reasoning Capabilities
|
|
openreview.net
|
Soft Mellowmax Monte Carlo Planning
|
|
openreview.net
|
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
|
|
openreview.net
|
Self-Improving VLM Judges Without Human Annotations
|
|
openreview.net
|
Duel-Evolve: Pairwise Preference Black-Box Optimization of LLM Responses
|
|
openreview.net
|
MimicAgent: Learning Quadruped Skills via Text-to-Trajectory Generation
|
|
openreview.net
|
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
|
|
openreview.net
|
CircuitBuilder: From Polynomials to Circuits via Reinforcement Learning
|
|
openreview.net
|
Generative Recursive Reasoning Models
|
|
openreview.net
|
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
|
|
openreview.net
|
Refining Large Language Models with Self-Generated Data Through Iterative Training
|
|
openreview.net
|
Inference-Time Scaling in Diffusion Models through Iterative Partial Refinement
|
|
openreview.net
|
In-Context Adaptation
|
|
openreview.net
|
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
|
|
openreview.net
|
Structure Enables Effective Self-Localization of Errors in LLMs
|
|
openreview.net
|
Self-Improvement via Fast Tree-search
|
|
openreview.net
|
Self-Adapting Agents for Automating Research Coding Workflows
|
|
openreview.net
|
Verifying the Verifiers: Failure Attribution for Agentic Benchmark Diagnostics and Training Data Curation
|
|
openreview.net
|
Just Enough Learning: GRPO-Guided Controllers for Hyperparameter Sweeps
|