
Closed
Posted
AI Agent Evaluator Remote · Contract · Flexible Hours We're hiring evaluators to help assess and benchmark cutting-edge AI agents. This is hands-on work at the frontier of agentic AI — not prompt writing, not basic QA. You will be designing complex real-world tasks, running them across multiple large language models, and producing the structured evaluations that shape how the next generation of AI agents are trained. What You'll Do You will design multi-stage agent tasks across real-world domains including health, finance, productivity, relationships, and personal exploration. Each task must require the agent to coordinate across multiple tools and systems, manage persistent state, handle realistic friction points, and produce a verifiable final artifact. Tasks that can be completed in a short, linear interaction are not complex enough for this project. You will run each task prompt across two AI models inside the OpenClaw environment, extract the full agent trajectories, and compare how each model planned, reasoned, used tools, and arrived at its final output. You will build binary evaluation rubrics that are atomic, objective, and self-contained — weighted on a scale from -5 to +5 — to measure task completion, instruction following, tool use, agent behavior, factuality, and safety. At least one negative-weight criterion is required per rubric set. You will annotate safety failures across seven domains including high-stakes actions, private data handling, ambiguous instructions, third-party prompt injections, and over-refusal. Each failure must be categorized by type, trajectory step, and action tier. You will write pytest-based unit tests in a [login to view URL] file that validate the agent's final system state using [login to view URL] as ground truth — confirming outcomes, not just intent. You will select the best-performing model trajectory, clone it, and iterate with the model until it passes all rubrics, producing a polished silver trajectory as the reference trace. You Must Be Able To Design complex multi-stage agentic tasks with real friction, constraint logic, and verifiable outcomes Write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models Annotate safety failures with the correct category from a defined failure taxonomy and assign the correct action tier Read and reason critically about full agent trajectories Write and validate basic Python unit tests using pytest Source every task idea from a real, publicly available online post about OpenClaw — no invented scenarios Strong Background In AI evaluation, rubric-based assessment, data annotation, agentic AI systems, AI safety concepts, prompt engineering, Python and pytest, and structured critical thinking. Important This is not a prompt-writing role or a basic chatbot quality assurance position. You are evaluating whether AI agents can architect solutions, coordinate tools, manage persistent state, and behave safely under real-world constraints. If a task can be solved without architectural reasoning or meaningful tool coordination, it does not meet the bar for this project.
Project ID: 40513036
25 proposals
Remote project
Active 1 day ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
25 freelancers are bidding on average $10 USD/hour for this job

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Matlab, Python, and similar tools. I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$25 USD in 40 days
7.1
7.1

Hi, I’d be excited to contribute as an AI Agent Evaluator. This role strongly matches my experience with agentic AI systems, LLM evaluation, structured rubric design, tool-using agents, Python-based validation, and safety-focused analysis. I have hands-on experience building and evaluating AI agents, RAG workflows, multi-step automation systems, tool-calling pipelines, and model reasoning flows. I understand that this is not basic prompt writing or chatbot QA; the goal is to test whether agents can plan, coordinate tools, handle persistent state, recover from friction, follow constraints, and produce verifiable outcomes. For this project, I can design realistic multi-stage agent tasks sourced from public OpenClaw-related posts, run them across multiple models, analyze full trajectories, and compare planning, tool use, reasoning quality, state management, factuality, and safety behavior. I can also create atomic binary rubrics with clear pass/fail criteria and meaningful weights from -5 to +5, including negative-weight criteria where needed. My background in AI engineering, agent workflows, Python, evaluation, prompt design, and structured critical thinking allows me to contribute carefully and consistently. I can work independently, follow detailed guidelines, and produce high-quality evaluation artifacts and polished silver trajectories.
$6 USD in 40 days
6.2
6.2

As an AI and Data Science specialist with a penchant for HW/SW development, Computer Vision, and Machine Learning (ML), I possess the exact skillset and experience needed to excel in your AI Agent Evaluator role. My previous work in ML enables me to read, comprehend and critically assess full agent trajectories to formulate detailed binary evaluation rubrics. This rigorous approach results in rubrics that reveal meaningful disparities between models and prompt their architectural development.
$5 USD in 40 days
4.6
4.6

With a solid foundation in programming, specializing in Python and pytest, I possess the skills needed to design multi-stage agentic tasks with verifiable outcomes involving AI systems. Critical thinking and discerning competencies acquired through years of experience will enable me to critically read and reason about complex agent trajectories. Furthermore, I am knowledgeable about AI safety concepts, adept at identifying potential safety vulnerabilities embedded within actions and meticulously categorizing them. My work as a web developer has necessitated in-depth knowledge of data monitoring, API integrations, and system responsiveness - skillsets that are instrumental to effective AI evaluation. Additionally, as an individual committed to delivering clean, maintainable code and all-encompassing tech solutions, my expertise dovetails well with your project requirements for structuring complex tasks with real-world friction points. To conclude, my resolve on producing quality deliverables in a timely fashion along with my propensity for long-term solutions aligns cohesively with the objective of your project. Let’s collaborate!
$8 USD in 40 days
4.0
4.0

Hi, I have experience with AI evaluation, LLM testing, rubric-based assessment, AI agents, Python, and pytest. I can design complex agentic tasks, evaluate model performance, create structured rubrics, analyze trajectories, and validate outcomes with reliable testing methodologies. Best regards, Shakila Naz
$5 USD in 40 days
3.8
3.8

Your project aligns perfectly with my expertise in designing and evaluating complex tasks for deep learning models. As a PhD candidate focused on high-performance computing and deep learning, I have experience in developing intricate data pipelines and benchmarking AI systems. My proposed price for this work is 20 USD/hr. This role requires meticulous attention to detail and the ability to design sophisticated evaluation frameworks, which are core skills I bring to the table. Looking forward to contributing to your cutting-edge project.
$20 USD in 40 days
2.0
2.0

I have strong Python/pytest skills, IBM AI Developer certification, and hands-on experience with LLM APIs and prompt engineering. I'm newer to formal agent evaluation rubrics but learn quickly and am genuinely interested in AI safety work. Happy to complete a trial task to demonstrate my evaluation quality.
$3 USD in 40 days
1.5
1.5

As an experienced full-stack developer with a deep understanding of AI engineering, my skillset optimally aligns with the objectives of this project. In my career, I've had the opportunity to explore and develop cutting-edge solutions using advanced AI technologies like LangChain, LangGraph, and RAG pipelines. These experiences have allowed me to hone my abilities in LLM API integrations as well as designed AI agents that connect external tools like what you were asking for the evaluation process. Additionally, I bring insightful technical writing skills to the table that can be instrumental in your project. Not only can I design complex multi-stage agentic tasks but also write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models - a key requirement for your project. Powered by my structured critical thinking skills, I've always been able to effectively analyze large datasets, and derive meaningful patterns. This experience will be valuable in reading and reasoning critically about full agent trajectories.
$5 USD in 40 days
1.6
1.6

I specialize in AI agent evaluation, rubric design, and structured benchmarking of LLM-based systems with a strong focus on reasoning, tool use, and real-world task complexity. I can design multi-stage agentic tasks, build atomic evaluation rubrics with clear scoring criteria, and accurately assess model trajectories using objective, verifiable outcomes. With experience in Python and pytest-based testing, I can contribute to high-quality, safety-aware evaluation pipelines that help improve next-generation AI agents.
$5 USD in 20 days
1.7
1.7

✨Hi. Hope you're doing well✨ This role is very well-aligned with my experience in AI systems evaluation, agent workflows, and structured benchmarking of LLM-based systems. I have experience workng on outlier and Alignerr. That's why this position is an ideal match with me. I have hands-on experience designing multi-stage agentic tasks, running model comparisons, and building structured evaluation rubrics with clear scoring logic. I’m comfortable working with trajectory analysis, tool-use evaluation, and identifying differences in planning, execution, and reasoning across models. I also have experience writing Python-based validation tests and working with structured evaluation pipelines where final system outputs are verified against expected results. I’m particularly interested in your focus on real-world friction, multi-step reasoning, and safety evaluation, as that matches how I’ve worked on similar AI assessment tasks. I’m available to start immediately and can adapt quickly to your OpenClaw workflow ?
$7 USD in 40 days
0.4
0.4

All 999 LLM routing attempts exhausted. Last error: Provider unavailable, no API key, or premium not allowed
$8 USD in 7 days
0.0
0.0

༺❖༻ Dear Client ༺❖༻ Thanks for posting about my specialist job area. Your required skills perfectly match my experience and work style. I have experience in AI evaluation, dataset annotation, prompt engineering, Python testing and building structured evaluation frameworks for LLM and agent workflows. I focus on objective scoring systems and reproducible testing pipelines. I can design multi stage agent tasks with real world constraints, build atomic rubrics with weighted scoring and evaluate model trajectories across planning, tool use, reasoning and final outputs. I also write pytest based verification scripts to validate final system states against ground truth data. I am comfortable analyzing long agent traces, identifying safety issues and producing structured evaluation reports that help improve model performance. I am also experienced working in agent style environments similar to OpenClaw where tool traces and structured outputs must be carefully validated. Given the opportunities to apply my passion and expertise to your project I am sure I will complete it perfectly on time and on budget. Let’s discuss your evaluation framework and tooling setup. Best regards, Glenn Bondoc
$5 USD in 40 days
0.0
0.0

Dear Client, Good morning . How are you? I hope this proposal finds you well. I'M A CERTIFIED & EXPERIENCED EXPERT, WELL VERSED WITH THE REQUIREMENTS FOR YOUR PROJECT "HIRING: Top-Tier AI Data Annotation & Evaluation Specialists -- 7." This is to inform you that I have KEENLY gone through your project description, CLEARLY understood all the project requirements as instructed in your project proposal and this is to let you know that I will perfectly deliver as desired. Being in possession of all stated required skills, (Technical Writing, Software Engineering, Data Visualization, Computer Vision, Excel, Data Science, Prompt Engineering, AI (Artificial Intelligence) HW/SW and Machine Learning (ML)), as this is my field of professional specialization having completed all certifications and developed adequate experience in the respective field, I hereby humbly request you to consider my bid for professional, quality and affordable services that meet all your requirements. I always guarantee timely delivery and unlimited revisions where necessary hence you are assured of utmost satisfaction when working with me. Please send me a message so that we can discuss more and seal the project. WELCOME.
$50 USD in 40 days
0.0
0.0

Hi employer, 1. Here’s what I can bring to your project: - Ability to design complex, real-world agent tasks with clear success criteria - Experience creating structured, atomic evaluation rubrics with measurable scoring logic - Strong attention to detail when comparing multi-step model trajectories - Comfort with Python and writing basic pytest-based validation tests - Analytical mindset for identifying reasoning, tool-use, and failure patterns 2. I can start immediately and work flexibly depending on your task pipeline. I’m also comfortable following strict evaluation frameworks and safety taxonomies as required. Would love to discuss your current evaluation setup and how I can plug into your workflow quickly. Best regards.
$5 USD in 20 days
0.0
0.0

Hello, Your project stood out because this is structured AI-agent evaluation rather than basic annotation or generic prompt writing. My background is in data science, Python, analytical validation, and technical documentation. I can support the work by: ● Designing realistic, multi-step tasks with clearly verifiable outcomes ● Reviewing agent trajectories and identifying where reasoning or tool use failed ● Writing atomic, weighted evaluation criteria ● Comparing outputs across multiple models consistently ● Classifying safety, reliability, and instruction-following failures ● Creating reproducible Python or pytest-based checks where appropriate ● Documenting results clearly enough for another evaluator to reproduce them I am methodical with ambiguous cases and would rather justify an evaluation against an explicit rubric than rely on subjective impressions. I can begin with a paid trial batch so that you can assess the precision of my evaluations before assigning a larger workload. Could you clarify whether the first assignment will focus mainly on task creation, trajectory evaluation, or writing validation tests? Best Regards, Sergej W.
$6 USD in 20 days
0.0
0.0

Hello, I am interested in the AI Agent Evaluator position. I have a Master's degree in Data Science and experience developing Python applications, APIs, and data-processing workflows. My background includes Python, SQL, Django, FastAPI, pandas, and structured analysis. I enjoy evaluating systems, designing reproducible processes, and documenting behavior clearly. I am particularly interested in agentic AI and the challenge of creating realistic multi-step tasks and objective evaluation criteria. I am comfortable working with Python and writing tests, and I value thorough, well-documented work. My experience includes building full-stack applications, working with databases, and analyzing data using Python. I am excited by the opportunity to contribute to AI evaluation and to help improve the next generation of AI agents. I am available on a flexible basis and can commit up to 10 hours per week. Thank you for your consideration.
$8 USD in 10 days
0.0
0.0

atlanta, United States
Payment method verified
Member since Oct 24, 2019
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$250-750 USD
₹1500-12500 INR
$2-8 USD / hour
₹750-1250 INR / hour
₹37500-75000 INR
$15-25 USD / hour
$15-25 USD / hour
£20-250 GBP
$3000-5000 USD
$20 USD
₹750-1250 INR / hour
₹600-1500 INR
₹12500-37500 INR
₹1500-12500 INR
₹1500-12500 INR
₹600-1500 INR
$250-750 AUD
₹12500-37500 INR
₹600-1500 INR
₹600-1500 INR