Join Project Spearmint, a multilingual AI response evaluation initiative that reviews large language model (LLM) outputs across different languages, with a focus on either Tone or Fluency. You must have native-level fluency in the target language and strong comprehension of English.
As an evaluator, you will review short, pre-segmented datasets and judge model-generated responses according to specific quality criteria. Your work will help validate evaluation frameworks and set baseline quality benchmarks for future model improvements.
Key responsibilities:
– Evaluate model responses in your native language, focusing on either Tone or Fluency.
– Assess the overall quality, accuracy, and naturalness of each response.
– Read the user prompt and two model replies, then score each using a five-point scale.
– Provide brief explanations for any extreme ratings.
Project overview:
Batch 1 – Tone: Check whether responses are helpful, insightful, engaging, and fair. Identify issues such as mismatched formality, condescension, bias, or other tonal problems.
Batch 2 – Fluency: Review grammatical correctness, clarity, coherence, and natural flow.
This is a project-based opportunity through CrowdGen, where you will join the CrowdGen Community as an Independent Contractor. If selected, CrowdGen will email you with instructions to create an account using your application email address. You will need to log in, reset your password, complete the setup requirements, and proceed with your application for the role.
Help shape the future of AI—apply today and contribute from home.
$3.70–$3.70 per hour
To apply for this job, please visit the application page
