The AgentHarness project is an open-source evaluation harness for Apodex-1.0, a verification-centric model for deep research. It provides a standard ReAct setup to reproduce public benchmark results. The project includes a quick start guide and supports various Apodex-1.0 variants.
The AgentHarness project can be used to evaluate the performance of Apodex-1.0 models on public deep-research benchmarks. It supports multiple benchmarks, including BrowseComp, BrowseComp-ZH, HLE-Text, and DeepSearchQA. The project provides a simple way to run smoke tests and configure environment variables.
The target audience for the AgentHarness project is researchers and developers interested in evaluating the performance of Apodex-1.0 models on deep-research benchmarks. The project requires some technical expertise, including knowledge of Python and deep learning frameworks.
The AgentHarness project can be monetized through consulting services, where experts help organizations set up and use the evaluation harness. Additionally, the project can be used to develop and sell pre-trained Apodex-1.0 models, or to offer cloud-based evaluation services. The project's maintainers can also offer support and maintenance services to organizations using the evaluation harness.