Data Scientist
Red Bank, NJ - USA
Job Summary
We are seeking a Data Scientist to build and curate the knowledge base for a generative AI/ML research effort and to develop the data and knowledge extraction pipelines for machine learning models. The role is responsible for extracting structuring and validating design knowledge into knowledge graphs and ontologies; for engineering the metadata annotation and provenance of program datasets; and for the analysis that turns results laboratory measurements and simulation output into actionable findings for the research team.
The ideal candidate is a strong applied data scientist with an interest in symbolic and structured representations of knowledge. Experience with formal methods and domain-specific languages (DSLs) is desired but not required.
- Design build and populate the research teams effort knowledge bases and maintain it under version control with provenance tracking
- Develop knowledge extraction pipelines (rule-based statistical and ML-assisted) that convert unstructured and semi-structured sources into structured queryable knowledge; establish quality metrics and validation procedures for extracted content
- Engineer the metadata annotation schema and packaging for program datasets (synthetic simulated and real collections) so that datasets are reproducible well documented and deliverable on the program data sharing schedule
- Perform exploratory and inferential analysis on training data simulation output and evaluation results; build dashboards and reports that show where generated waveforms succeed or fail against objectives and feed findings back to AI model researchers and engineers
- Contribute data and analysis sections to design reviews monthly status reports and dataset documentation; coordinate with academic subcontractors on shared data and knowledge resources
Required Qualifications:
- Bachelors Degree or higher in Computer Science Statistics Mathematics Electrical Engineering or a related technical field
- 5 years of applied data science or data engineering experience (or MS with 3 years) with a record of delivering data pipelines and analyses that other engineers and researchers depend on
- Strong Python proficiency including the scientific stack (NumPy pandas SciPy scikit-learn) and experience with at least one deep learning framework (PyTorch preferred)
- Hands-on experience with knowledge graphs ontologies or structured knowledge representation (RDF/OWL property graphs such as Neo4j or equivalent) and with querying and validating them (SPARQL Cypher SHACL or similar)
- Experience designing data schemas metadata standards and annotation workflows for large scientific or engineering datasets including data versioning and provenance
- Solid statistical foundations: experimental design hypothesis testing uncertainty quantification and the ability to explain results to technical and non-technical audiences
- Experience with SQL and with at least one workflow or pipeline orchestration tool (Airflow Prefect Dagster DVC or equivalent)
- Ability to produce clear documentation including data dictionaries dataset cards and analysis reports
- US Citizenship
Desired Qualifications:
- Experience with formal methods or formal verification such as SMT solvers (Z3 cvc5) model checkers property-based testing (Hypothesis) or proof assistants particularly applied to validating generated programs or signal processing pipelines
- Experience designing or implementing domain-specific languages: grammar design parser generators (ANTLR Lark or similar) type systems intermediate representations or compiling DSL programs to executable code
- Background in symbolic AI or neuro-symbolic methods: logic programming constraint solving rule engines or program synthesis
- Familiarity with digital signal processing and communications fundamentals (modulation filtering coding channel effects) or with GNU Radio and software-defined radio data formats (I/Q sample handling SigMF or similar metadata standards)
- Prior work on IARPA DARPA or similar government research programs including data sharing plans privacy protection plans and delivery of datasets to independent T&E teams
- Experience building causal or probabilistic models from structured knowledge or working alongside causal inference researchers
- MS or PhD in Computer Science Statistics Electrical Engineering or a related technical field
- Willingness and ability to obtain Secret security clearance
Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the worlds leading mission capability integrator and transformative enterprise IT provider we deliver trusted highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land sea space air and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day our employees do the cant be done by solving the most daunting challenges facing our customers. Visit to learn how were keeping people around the world safe and secure.
Required Experience:
IC
About Company
Peraton provides innovative solutions for the most sensitive and critical programs in government today, developed and executed by scientists, engineers, and other experts.