Microsoft

Pre-Training Text Data

United States, Multiple Locations, Multiple Locations
✓ Verified live on the employer's own system · added 193 days ago
Save search
Senior · 12+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 12+ years

Skills & tools

Data AnalysisProgrammingPython
Apply on company site ↗ See your fit → free

Full job description

Develop novel data collection strategies Improve dataset quality and integrity Create high-quality datasets for training and evaluation; run experiments on new datasets (data ablations) to assess their impact and determine the most effective data.

Develop and maintain scalable data pipelines for text data ingestion, preprocessing, filtering, and annotation. Analyze real-world text datasets to assess quality, diversity, relevance, and identify areas for improvement.

Build lightweight tools and workflows for dataset auditing, visualization, and versioning. Collaborate with Safety, Ethics, and Governance teams to ensure datasets meet standards for quality, privacy, and responsible AI practices.

Embody our culture and values. Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or related technical discipline AND 4+ years technical engineering experience with coding in languages including, but not limited to, Python and common data libraries (Pandas, NumPy, etc.) Master's Degree in in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or related technical discipline AND 8+ years technical engineering experience with coding in languages including, but not limited to, Python and common data libraries (Pandas, NumPy, etc.) OR Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or related technical discipline AND 12+ years technical engineering experience with coding in languages including, but not limited to, Python and common data libraries (Pandas, NumPy, etc.) OR equivalent experience.

2+ years of experience in data analysis or data engineering, including work with large-scale datasets that are unstructured or semi-structured.

Proficiency in statistics and exploratory data analysis methods. Familiarity with data processing frameworks such as Spark, Ray, or Apache Beam.

Ability to communicate technical findings clearly to research and product teams.

More jobs at Microsoft

Similar jobs near United States, Multiple Locations, Multiple Locations

Tell me when more Principal Cloud Solution Architect, AI Data, Partner Solutions jobs post near United States, Multiple Locations, Multiple Locations We re-check every listing against the employer’s own board — no résumé needed.

Search Pre-Training Text Data jobs near United States, Multiple Locations, Multiple Locations → Browse all live jobs

This posting was published by Microsoft on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.