AI Evaluators: Assessing A Shopping Assistant

Terac - United States - original posting ->
Status
Open
Remote policy
Remote
Employment type
Not stated
Salary
Not stated
Categories
AI-Evaluator, AI-Product-Evaluator, AI-Chatbot-Evaluator, AI-Agent-Evaluation, AI-Evaluation, AI-Quality-Evaluator, Chatbot-Conversational-AI-Evaluator, AI-Product-Evaluation, AI-Assessor, AI-Chatbot-Evaluation
Tech
remote-country
Source
himalayas
First observed
2026-09-03 02:36 UTC
Last seen
2026-09-03 02:36 UTC
Source claims posted
2026-09-03 02:09 UTC
Consecutive misses
1 of 10

What the posting says

What We're Researching

We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.

How It Works

You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.

Who This Is For

This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.

What You'll Do

Review real user interaction traces with an AI shopping assistant

Identify logical failures, inaccuracies, or poor recommendations in the text

Create structured rubrics and verifiers to judge response quality

Commit to 20+ hours per week of evaluation work on our internal platform

Who Should Apply

Experience in data evaluation, quality assurance, or AI training

Strong analytical skills with the ability to spot subtle errors in text

Familiarity with e-commerce search and digital shopping experiences

Ability to commit to a sustained workload of 20+ hours per week

Compensation

$50 per hour

Ready to participate?

Start your paid interview now

About Terac

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

Learn more at or on YouTube at @jointerac.

Originally posted on Himalayas

Quality

Completeness: 65%

Not enough history yet to judge honesty signals.

Timeline

  1. *
    #541894 2026-09-03 02:36 UTC
    Published
  2. o
    #542565 2026-09-03 04:39 UTC
    Not seen
    Miss 1 in a row