Appier said its research paper, Joint Optimization of Tool Creation and Use for Large Language Model Agents, has been accepted at NeurIPS. The paper introduces SMITH, which the company described as a reinforcement learning framework that lets AI build tools and use them in a single training loop, with each tool refined as it is applied to problems. The company said the work addresses a challenge in agentic AI: models can create tools but often struggle to use them well.
According to the release, many AI systems still depend on engineers to build APIs or configure fixed tools in advance, and those tools must be rebuilt when data sources, tasks, or business needs change. Even when AI creates tools, existing methods usually assign creation and use to separate models, so the tool-building model receives little feedback on real-world performance and cannot easily tell whether a tool is clearly described, works reliably, or can be called correctly by other models. SMITH places both skills in one training loop, the company said, so results feed back when a tool has a vague description, poorly designed parameters, or fails to run. Chieh-Yen Lin, a research scientist at Appier, said that during training the model sees only a tool's description and parameter specifications, not its underlying code, which makes the clarity of descriptions and whether a tool can be called correctly direct feedback.
The research found that a model of about 4 billion parameters trained with SMITH built tools that outperformed those from every other method in the study on unseen tasks, and beat a baseline in which a roughly 30-billion-parameter model built tools on the fly. Tools built by the small model handled new tasks when used by a model of about 350 million parameters and also improved larger models. SMITH trains from 4 simple examples and tests on 16 harder, previously unseen problems, keeping only tools that solve new problems and adding them to a shared library. Average output fell from 3,206 tokens with conventional step-by-step reasoning to about 100 tokens, which the release describes as roughly a 32-fold gain in efficiency.
Dr. Chih-Han Yu, CEO and co-founder of Appier, said humans turn their problem-solving experience into tools so they never have to start from scratch, and that AI agents are evolving in the same way. He said the research shows agents can learn to build tools, continuously refine them, and share proven tools across models of all sizes, making multi-agent collaboration more efficient and scalable. Appier said many daily operations are repetitive, from converting financial metrics and processing data to querying reports, checking rules, and routing customer service cases, and that AI could turn these methods into shared tools that any agent can call.