Join Project Spearmint, a multilingual AI response evaluation initiative that reviews large language model (LLM) outputs across different languages, focusing on either Tone or Fluency. You must be a native-level speaker of the target language and have strong English comprehension.
As an evaluator, you will review short, pre-segmented datasets and rate model-generated responses based on specific quality criteria. Your work will help validate evaluation frameworks and establish baseline quality benchmarks for future model improvements.
Key responsibilities:
– Evaluate model responses in your native language, focusing on either Tone or Fluency.
– Judge the overall quality, accuracy, and naturalness of each response.
– Read the user prompt and two model replies, then score each on a five-point scale.
– Provide brief explanations for any unusually high or low ratings.
Project overview:
Batch 1 – Tone: Determine whether responses are helpful, insightful, engaging, and fair. Flag issues such as inappropriate formality, condescension, bias, or other tonal problems.
Batch 2 – Fluency: Check grammatical correctness, clarity, coherence, and how naturally the response flows.
This is a project-based opportunity through CrowdGen. If you’re selected, CrowdGen will email you instructions to create an account using your application email address. You will need to log in, reset your password, complete the setup requirements, and then proceed with your application for the role.
Help shape the future of AI—apply today and contribute from home.
$3.70–$3.70 per hour
To apply for this job, please visit the application page
