Data Annotation Specialist, Generalist

Cohere · Canada
contract junior Posted 23 hours ago
Apply Now Stand out: build a proof-of-work pitch →

Free GitHub-based preview. Direct apply stays one click away.

Get weekly job alerts like this →

Hiring for this role?

AI Market Demand Pack · $29 one-time

Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.

See the live sample →
llm

About this role

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company co-headquartered in Toronto and San Francisco, with key offices in London, New York City, Montreal, Seoul, Germany and Paris. Join us! Why this role? We are on a mission to build machines that understand the world and make them safely accessible to all. Data quality is foundational to this process. Machines (or Large Language Models, to be exact) learn in similar ways to humans, by way of feedback. By labelling, ranking, auditing, and correcting model output, you will improve Large Language Models' performance for iterations to come, thus having a lasting impact on Cohere's technology. We are hiring Generalist professionals with broad backgrounds that span multiple consumer-facing or personal domains. This is a judgment-driven role, not passive data entry. You will review, assess, and provide structured feedback across a broad and evolving range of tasks, evaluating, stress-testing, and improving our models on English-language data spanning multiple modalities (text, image, and structured formats such as JSON, CSV/TSV, and Markdown). This is a great opportunity for professionals with strong analytical skills to contribute to high-impact annotation projects. Please Note: This is a part-time independent contractor position available within Canada only. We seek candidates who can commit to 16 hours per week at a CAD $30/hour contract rate. This role is BYOD 💻 (Bring Your Own Device, laptop). This position is remote. As an Data Annotation Specialist, you will: - Evaluate and rank model outputs: Complete preference and comparison tasks, assessing which responses best conform to project guidelines for accuracy, helpfulness, tone, and safety, and writing clear justifications for your judgments. - Stress-test and break models: Probe models adversarially to surface failure modes, unsafe behavior, and capability gaps, and document reproducible cases that engineering and research teams can act on. - Create datasets: Author high-quality prompts, responses, and exemplars to build training and evaluation datasets, following detailed specifications and editing machine-written or human-written outputs to standard. - Build and apply rubrics and taxonomies: Contribute to the design of grading criteria and rubrics, then apply them consistently to produce structured, high-quality annotations across task types. - Annotate and correct multimodal data: Label, audit, and rectify inaccuracies across text, image, and structured data, maintaining a high standard of data integrity and accuracy. - Calibrate and maintain consistency: Participate in calibration exercises and inter-annotator agreement checks to align on standards, and flag ambiguous or uncovered edge cases rather than guessing, since a single misjudgment replicated at scale degrades a model. - Adapt to experimental work: Take on new and evolving task types as project needs shift, applying sound judgment in areas where guidelines are still being developed. - Report on model performance: Surface and communicate quality and performance trends in model and agent behavior, giving cross-functional partners clear, well-evidenced feedback on where models succeed, fail, and degrade. You may be a good fit if you have: - 1+ years of experience in AI data annotation, LLM evaluation, content moderation, research, or a related analytical role, with exposure to quality assurance, and/or preference ranking. - Experience applying detailed guidelines to complex and often ambiguous content, with strong contextual and sociocultural judgment, sensitivity to nuance, tone, and register, and the ability to reason well in cases where there is no single correct answer. - Comfort with ambiguity: a willingness to flag unclear or uncovered edge cases rather than guess, and to work productively on novel, experimental tasks whose definitions are still evolving. - A sharp, curious eye for inconsistencies, subtle errors, and model failure modes, including the instinct to probe models adversarially and surface where they break. Familiarity with how large language models behave, such as hallucination, sycophancy, and instruction-following gaps, is a plus. - Excellent command of written English and strong reading comprehension, with the ability to clearl

Similar Jobs

Related searches:

Remote Jobs Junior Jobs Remote Junior Jobs Junior Machine LearningJunior NLP & Language AI llm

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.