TikTok is a short-form mobile video platform whose Feed Safety-Model & Data Intelligence Team ensures that AI models meet rigorous standards before making content safety decisions. The intern will design AI safety model evaluations, build ground truth datasets, analyze model performance, collaborate with algorithm engineers, communicate governance standards, and improve scalable evaluation processes.
Responsibilities
Design and execute evaluation plans for AI safety models — define evaluation objectives, select appropriate metrics, and determine what "good" looks like for each use case
Build and maintain high-quality ground truth datasets — design data sampling strategies, develop data cleaning pipelines, and ensure labeling consistency and accuracy
Analyze model performance using statistical methods (sampling design, confidence intervals, error analysis) to produce actionable insights for algorithm teams and stakeholders
Collaborate with algorithm engineers to translate evaluation findings into concrete model improvement directions; participate in prompt design and model configuration iteration
Communicate evaluation results and governance standards to cross-functional partners (Policy, Operations, Algorithm); align on definitions and help calibrate quality expectations
Continuously improve evaluation processes — identify gaps, propose methodology upgrades, and ensure our evaluation systems scale with model and policy evolution
Qualification
Required
Currently pursuing an Undergraduate/Master's in Statistics, Computer Science, Data Science, Public Policy, or closely related quantitative fields
Solid grasp of applied statistics — sampling, hypothesis testing, confidence intervals, distribution analysis — and ability to apply these to real measurement problems
Foundational understanding of AI/ML concepts (classification, NLP, LLMs, precision/recall); comfortable discussing model behavior with engineers
Interest in governance, policy, or content safety; appreciation for the complexity of defining 'right' and 'wrong' at scale
Strong structured thinking — able to decompose ambiguous problems into clear goals and prioritized actions; goal-oriented and hypothesis-driven
Ability to communicate in both English and Chinese to collaborate with global and China-based stakeholders
Preferred
Internship or project experience in Trust & Safety, AI/ML product, model evaluation, or policy-related work
Hands-on experience with prompt engineering, LLM-based evaluation, or building evaluation datasets
Familiarity with safety-specific challenges: adversarial content, inter-annotator disagreement, or human-in-the-loop systems
Coursework or research in AI governance, computational social science, or interdisciplinary areas combining technology and policy
Experience with experimental design and statistical modeling beyond introductory level
Benefits
Hands-on experience and industry exposure
Opportunities to apply knowledge to real-world challenges
Opportunities for personal and professional growth
Practical experience and opportunities to explore potential career paths
Participation in social events, learning programs, and development workshops alongside industry professionals
Day one access to health insurance, life insurance, wellbeing benefits and more
10 paid holidays per year
Paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year)
Interns who are not working 100% remote may also be eligible for housing allowance
TikTok is a short-form video entertainment app and social network platform. It is a sub-organization of ByteDance.