Data Engineer
McLean, MD - USA
Department:
Job Summary
For over 25 years Edgesource Corporation has served as an innovative technology service provider for the Department of Defense (DOD) Department of Homeland Security (DHS) Department of State (DOS) the U.S. Intelligence Community Law Enforcement and other federal state and commercial clients locally nationally and abroad. From providing boutique technical solutions in support of the DOD Counter Unmanned Aerial Systems (CUAS) mission set to addressing the most critical Cybersecurity threats facing our nation as a prime contractor with the DHS Cybersecurity & Infrastructure Security Agency (CISA) a career at Edgesource is an opportunity to do meaningful interesting and impactful work.
We are seeking a Data Engineer to work with a small team to build complex data flows for a custom application. Successful candidate will have advanced Python programming skills familiarity with Java an understanding of data security privacy governance and compliance principles and a demonstrated history of building production data pipelines and ETL workflows at scale.
Building end-to-end data pipelines leveraging Python
Using orchestration tools to deploy data pipelines including configuring and updating Spark Jobs
Containerizing and deploying applications in cloud environments like AWS.
Working with MySQL and PostgreSQL including performance tuning schema design and query optimization for complex analytical workloads.
Leveraging industry standard tools for code control (Git IaaC control etc.)
Working with data catalogs tracking data lineage and handling a variety of data formats including Geospatial.
Using Bash scripting for automation and data processing tasks
Integrating Al/ML services and models
Work with stakeholders to understand data requirements assess feasibility and design appropriate solutions with minimal oversight
Leverage strong problem-solving and debugging skills for data quality issues pipeline failures and performance bottlenecks
Leverage a background in large-scale data migration or platform modernization efforts
Contribute to data engineering documentation best practices and design patterns.
Active TS/SCI Full Scope Polygraph
Minimum of 5 years experience
Demonstrated experience building production data pipelines and ETL/EL workflows at scale
Proficiency with Apache Spark and PySpark for distributed data processing
Advanced Python programming skills including data manipulation libraries (Pandas NumPy) and data engineering best practices
Understanding of data security privacy governance and compliance principles
Experience with workflow orchestration tools (such as Step Functions Airflow)
Familiarity with containerization (such as Docker or Podman) and deploying data applications in cloud environments
Experience with AWS services (S3 Lambda Step Functions)
Experience with PostgreSQL and MySQL in production environments including performance tuning and schema design
Demonstrated experience with SQL and query optimization for complex analytical workloads
Experience with version control (Git) and Cl/CD practices for data pipelines
Demonstrated ability to work with stakeholders to understand data requirements assess feasibility and design appropriate solutions with minimal oversight
Strong problem-solving and debugging skills for data quality issues pipeline failures and performance bottlenecks
Experience with data Lakehouse architecture using Apache Iceberg
Hands-on experience configuring deploying and integrating data platform components:
Apache Ranger (access control and data governance)
Trino (distributed SQL query engine)
Data catalogs (Unity Catalog OSS Apache Polaris etc.)
Apache Superset data visualization and dashboarding)
Proficiency with Bash scripting for automation and data processing tasks
Experience with Infrastructure as Code (Terraform or CloudFormation) for data infrastructure
Familiarity or experience with tracking data lineage and associated tooling such as Open lineage
Familiarity or experience with Java
Familiarity with data quality frameworks testing methodologies and validation strategies
Background with large-scale data migrations or platform modernization efforts
Experience integrating Al/ML services and models (translation OCR speech-to-text NLP language detection topic modeling) LLMs and RAG retrieval-augmented generation) pipelines
Familiarity with geospatial data processing H3 PostGIS or similar)
Contributions to data engineering documentation best practices and design patterns
Experience with NoSQL databases (DynamoDB etc.)
As an ISO 9001:2015 certified and CMMI Level 3 appraised small business Edgesource specializes in providing a variety of technical solutions to include software development database services enterprise networking data center virtualization and management support. We are always seeking top-talent to join our team in helping to address the most critical technical challenges facing our nation.
At Edgesource we understand that our employees are our greatest asset and as such we offer a wide array of benefits to support the well-being of our staff to include:
Flexible PTO Policy 11 Paid Holidays
Flexible Work Schedules (Remote / Hybrid)
Medical / Dental / Vision / Flexible Spending Account (FSA)
401k Plan with Match
Tuition & Professional Development Support
Commuter Benefits
Bonus & Employee Referral Programs
Career Growth Opportunities
All qualified applicants will receive consideration for employment without regard to race color religion sex disability age sexual orientation gender identity national origin veteran status or genetic information. Edgesource is committed to providing access equal opportunity and reasonable accommodation for individuals with disabilities in employment its services programs and activities. To request reasonable accommodation please contact our Recruiting Department by email at or by phone at
Required Experience:
IC