Germany-Based English
Vor 2 Tagen
München, Brandenburg, Deutschland
Rex.zone
Remote
Vollzeit
30 $ - 50 $ Festanstellung
Kostenlos per E-Mail oder Google
Speichern Sie diesen Job und halten Sie Ihre Suche organisiert
Erstellen Sie ein kostenloses Konto, um Jobs zu speichern, Benachrichtigungen zu erstellen und von Ihrem Dashboard aus zu diesem Eintrag zurückzukehren.
Kostenlos per E-Mail oder Google
Wenn Sie fortfahren, akzeptieren Sie unsere Nutzungsbedingungen & Datenschutzerklärung.
Overview
Rex.zone is hiring Germany-based English/German AI Generalist Trainers to support RLHF and large language model evaluation by assessing, ranking, and quality-checking model outputs to strengthen training data quality and drive model performance improvement.
About The Role
In this role, you will evaluate and improve AI/LLM behaviors through RLHF-style workflows. You will review prompts and model responses, validate factuality and instruction-following, apply annotation guidelines compliance standards, and write clear rationales that improve training data quality.
Key Responsibilities
• Perform large language model evaluation across English and German prompts and outputs
• Rank multiple model-generated responses using defined rubrics and consistent reasoning
• Conduct QA evaluation checks by auditing items for annotation guidelines compliance
• Write concise rationales that justify rankings and identify reasoning errors
• Validate outputs for safety, policy adherence, and content safety labeling needs
• Identify edge cases and failure modes; document findings for model performance improvement
• Apply data labeling standards to create reliable feedback signals for RLHF pipelines
• Track discrepancies, escalate guideline questions, and support consistent quality targets Basic Qualifications
• Must be based in Germany and able to work remotely from Germany
• Fluency in English and German (strong reading and writing skills in both)
• Strong analytical skills; ability to evaluate arguments, logic, and evidence
• Exceptional attention to detail and consistency in applying rubrics
• Comfortable working with sensitive or policy-relevant content under safety labeling rules Preferred Qualifications
• Experience with AI data labeling, RLHF, prompt evaluation, or LLM evaluation
• Familiarity with LLM failure patterns (hallucinations, instruction drift, bias, safety issues)
• Background in linguistics, translation, writing/editing, QA, research, or content moderation Compensation Hourly pay range: $30–$50 USD per hour (base), based on skills, language proficiency, and evaluation quality.
How To Apply
Please apply with an updated resume and a brief note describing your English/German proficiency and any experience with evaluation, QA, annotation, or AI/LLM workflows.
About The Role
In this role, you will evaluate and improve AI/LLM behaviors through RLHF-style workflows. You will review prompts and model responses, validate factuality and instruction-following, apply annotation guidelines compliance standards, and write clear rationales that improve training data quality.
Key Responsibilities
• Perform large language model evaluation across English and German prompts and outputs
• Rank multiple model-generated responses using defined rubrics and consistent reasoning
• Conduct QA evaluation checks by auditing items for annotation guidelines compliance
• Write concise rationales that justify rankings and identify reasoning errors
• Validate outputs for safety, policy adherence, and content safety labeling needs
• Identify edge cases and failure modes; document findings for model performance improvement
• Apply data labeling standards to create reliable feedback signals for RLHF pipelines
• Track discrepancies, escalate guideline questions, and support consistent quality targets Basic Qualifications
• Must be based in Germany and able to work remotely from Germany
• Fluency in English and German (strong reading and writing skills in both)
• Strong analytical skills; ability to evaluate arguments, logic, and evidence
• Exceptional attention to detail and consistency in applying rubrics
• Comfortable working with sensitive or policy-relevant content under safety labeling rules Preferred Qualifications
• Experience with AI data labeling, RLHF, prompt evaluation, or LLM evaluation
• Familiarity with LLM failure patterns (hallucinations, instruction drift, bias, safety issues)
• Background in linguistics, translation, writing/editing, QA, research, or content moderation Compensation Hourly pay range: $30–$50 USD per hour (base), based on skills, language proficiency, and evaluation quality.
How To Apply
Please apply with an updated resume and a brief note describing your English/German proficiency and any experience with evaluation, QA, annotation, or AI/LLM workflows.