I am unable to reproduce MCQ 71.35 on the DriveLMM-o1 benchmark.
Could you please provide a process for reproducing this using the current code on GitHub and the agent_think model you provided?
Or is it that the current code and model are insufficient for reproduction? If so, could you provide a model and code that can reproduce the problem?
I followed the standard pipeline: inference_withtool.sh -> gen_tool_result_data.py -> prepare_data.py -> inference_agentthink.py.
I am unable to reproduce MCQ 71.35 on the DriveLMM-o1 benchmark.
Could you please provide a process for reproducing this using the current code on GitHub and the agent_think model you provided?
Or is it that the current code and model are insufficient for reproduction? If so, could you provide a model and code that can reproduce the problem?
I followed the standard pipeline:
inference_withtool.sh->gen_tool_result_data.py->prepare_data.py->inference_agentthink.py.