
### DataSet Evaluation

| Metric                        | Value |
|-------------------------------|-------|
| execution_match               | 0.3269230769230769 |
| execution_match_non_empty     | 0.39473684210526316 |
| total                         | 104 |
| total_non_empty               | 76 |
| total_gt_non_empty            | 100 |
| correct                       | 34 |
| correct_non_empty             | 30 |
| count_gt_read_failure         | 4 |
| count_response_read_failure   | 27 |

