PerkinElmer is seeking a Data Science Intern to support its Asset Intelligence Service, which normalizes, classifies, and enriches customer asset records for laboratory instruments. The intern will learn and document the service, build measurement standards, analyze classification accuracy and attribute coverage, and own a defined improvement workstream involving data matching and text classification.
Trace an asset record from raw customer intake through normalization, classification, and enrichment to the form it takes in the platform, and be able to explain each step to someone who has never seen it
Reproduce the pipeline on a sample set and confirm you get the same answers the production service does
Write down where the service makes a judgment call, what evidence it uses, and which of those calls are the fragile ones. This document does not exist today and the team needs it
Build and maintain a gold-standard reference set of correctly classified instruments, which is the thing the service is currently missing and the thing every accuracy claim depends on
Score match rate and classification accuracy against that set, reported by equipment class rather than as one headline number, so the weak classes are visible instead of averaged away
Re-score after every change, so improvement is demonstrated rather than asserted
Cluster near-duplicate records and reconcile them as a group with their alias set, rather than one record at a time. Records that differ only by punctuation or spacing are the case the current one-at-a-time approach cannot resolve
Work on the text that feeds the classifier, separating technical capability language from application and use-case language. Scraped marketing copy about what an instrument has been used for is a known source of misclassification
Analyze attribute coverage across the record base, field by field, so the team knows which attributes are populated well enough to build on and which are not before anything gets promised to a customer
Keep the code and queries in the team's repository (notebooks are fine) in a state someone else can pick up and run
Write short method notes alongside the code. Much of the value of this role is in the write-up, because the point is that the team can repeat and extend the work after the term ends
Take on ad hoc data pulls and analysis the team needs at short notice. Everyone on a team this size carries some of that, and it is also the fastest way to see how the platform is actually used
Qualification
Required
Currently enrolled in a bachelor's degree program in data science, statistics, computer science, or a related field
Working knowledge of Python and SQL
Available approximately ten hours per week during the academic term
Preferred
Coursework or project experience in data cleaning, record matching, deduplication, or classification
Familiarity with pandas and Jupyter, or the equivalent in R
Some exposure to AI-assisted text processing, along with the instinct to verify what it returns rather than accept it. Knowing how a plausible wrong answer happens matters more here than knowing how to prompt
Comfortable working remotely from written direction with limited supervision, and inclined to ask early when something is ambiguous instead of guessing and proceeding
Clear written communication. The analysis is only worth what the write-up conveys
Interest in scientific instrumentation or laboratory operations. No prior domain knowledge is assumed and none is required to start
Benefits
PerkinElmer focused on improving the health and safety of people and the environment.