Research questionHow can we train and evaluate LLM agents for tool use across single- and multi-turn workflows with serial or parallel calls?Tool-use studies often represent function-call interactions differently and cover uneven trajectory structures. This makes it difficult to align training data with benchmarks or compare agents when workflows vary in turn depth and call execution.