How to Determine if a Tool Has Truly Executed Successfully?
Check the exit status, actual output, and generated artifacts simultaneously.
Don't Rely Solely on a Single Success Summary
The Agent's natural language summary needs to correspond with the tool's output. Command start success, program exit success, and task goal achievement are three different levels: a program may exit normally but still not process the specified file.
First verify the target environment and input, then check whether the output sufficiently supports the final conclusion.
Judge by Visible Evidence
For example, in a report generation task, check that the file exists and the content is complete; for a program check task, look at the actual results and applicable conditions, not just whether there are red error messages.
For example, if the report states "check passed," ask which files were covered by the check, what inputs were used, and whether any items were skipped. Correct exit status only supports judgment at the execution level; whether your goal was achieved still needs to correspond with the acceptance criteria agreed upon at the start.
- Check the command execution status and key outputs.
- Confirm that the processed files, versions, or targets are as expected.
- Inspect the generated files and required fields.
- Clearly state any uncovered scope requirements.
When Output Is Very Long or the History Session Has Ended
The tool may retain limited output or provide a save location; do not assume the conversation preview shows the entire log. Have the Agent extract relevant segments and preserve the original results. "Running" in historical execution records does not necessarily prove the process is still running at this moment; when resuming a task, first verify the current status and artifacts before deciding the next step.