Research questionHow can cold-start LLM agents reliably improve an initial procedural skill?An initial skill may be syntactically valid yet fail during execution, while expert authoring is costly and one-shot generation may not reflect how agents actually perform tasks. The central difficulty is improving the skill when accumulated self-evolution trajectories are unavailable.