Engineer III GT Storage and Backup
Job Summary
Position Summary
The Engineer III is an experienced infrastructure engineer within Global Technology who designs implements automates secures and supports enterprise storage and data protection services. The role combines deep technical ownership with engineering practices moving beyond traditional administration toward platform reliability cyber resilience infrastructure as code observability and measurable business outcomes.
This engineer independently delivers solutions of significant size and complexity across storage SAN backup recovery and hybrid-cloud services. The role contributes to architecture and technical decisions supports globally distributed production environments mentors other engineers and translates business requirements into scalable recoverable secure and cost-effective capabilities.
Essential Functions and Core Competencies
Technical Depth and Delivery
- Design implement test upgrade and support enterprise block file object SAN and backup platforms across on-premises and hybrid-cloud environments.
- Own complex engineering work from requirements and design through change implementation validation documentation operational handover and lifecycle support.
- Configure and optimize storage provisioning replication snapshots performance tiering deduplication compression retention backup policies and recovery workflows.
- Diagnose cross-domain incidents involving arrays fabrics hosts virtualization databases applications backup media network dependencies and cloud services.
- Perform peer reviews for scripts automation configurations designs recovery runbooks and implementation plans; provide actionable technical feedback.
- Participate in on-call support major incident response root-cause analysis problem management and corrective-action follow-through.
Architectural Scale
- Design solutions of significant size and complexity that meet availability performance security compliance capacity RPO and RTO requirements.
- Contribute to architecture reviews standards reference patterns lifecycle roadmaps and technology-selection decisions for storage and data protection.
- Evaluate dependencies failure domains replication topology data locality licensing cost operational supportability and vendor lock-in when recommending solutions.
- Engineer high availability disaster recovery cyber recovery and business-continuity capabilities including isolated or immutable copies and repeatable recovery validation.
- Plan and execute platform upgrades migrations refreshes consolidations decommissions and data movement with controlled risk and rollback strategies.
Automation Platform Engineering
- Build reusable automation using PowerShell Python REST APIs vendor SDKs and orchestration tools such as Ansible or AWX.
- Manage scripts infrastructure definitions configuration and documentation through Git-based source control peer review branching and versioning practices.
- Integrate infrastructure automation with CI/CD workflows such as Azure DevOps or GitHub Actions including validation approvals secrets handling and auditable deployment evidence.
- Develop self-service workflows and policy-driven provisioning that reduce manual tickets improve consistency and enforces guardrails.
- Apply secure engineering practices including least privilege service-account governance credential rotation certificate lifecycle management vulnerability remediation and separation of duties.
- Use AI-assisted engineering responsibly for analysis code generation documentation and operational insights while validating outputs and protecting confidential data.
Reliability Observability and Operational Excellence
- Define and monitor service health indicators for availability latency throughput capacity backup success recovery readiness replication and platform risk.
- Implement actionable monitoring alerting dashboards event correlation and ticket integration using enterprise observability and ITSM platforms.
- Reduce alert noise and recurring incidents through automation problem elimination error budgets runbook improvement and preventive maintenance.
- Perform trend analysis and capacity forecasting; translate consumption growth performance and support lifecycle into timely CAPEX/OPEX recommendations.
- Maintain accurate configuration dependency risk operational and recovery documentation; create training material and reusable knowledge articles.
- Measure improvements through outcomes such as provisioning lead time automation coverage ticket reduction change success restore success capacity efficiency and customer satisfaction.
Cyber Resilience Governance and Recovery Assurance
- Implement layered protection using immutability isolation encryption multifactor authentication role-based access audit logging retention controls and secure administrative paths.
- Partner with cybersecurity risk compliance application database and business teams to define protection tiers recovery priorities and evidence requirements.
- Conduct scheduled restore tests application-consistent recovery tests disaster-recovery exercises and cyber-recovery simulations; document results and drive remediation.
- Assess unprotected workloads policy drift privileged access ransomware exposure end-of-support risk and recovery gaps; maintain prioritized remediation plans.
- Ensure changes and operations comply with organizational standards change controls data handling requirements and applicable legal or regulatory obligations.
Collaboration Influence and Business Impact
- Collaborate across engineering teams and with product security network compute cloud database application facilities and service-management stakeholders.
- Communicate technical options operational risks costs dependencies and trade-offs to technical and non-technical audiences.
- Influence technical outcomes through evidence-based recommendations design reviews standards and constructive consensus-building.
- Mentor less experienced engineers lead knowledge-sharing sessions and improve team capability through reviews pairing documentation and coaching.
- Coordinate with vendors and service providers on architecture support cases escalations maintenance licensing roadmap discussions and professional services.
- Apply business context to prioritize work that improves customer experience resilience operational effectiveness and value for the organization.
Storage and SAN Engineering
- Engineer and support enterprise storage platforms such as Pure Storage Dell PowerStore Hitachi VSP/HNAS and other approved block file NAS or object technologies.
- Administer Fiber Channel fabrics and management platforms such as Brocade switches and SANnav including zoning masking path resiliency firmware upgrades health checks and single-path remediation.
- Perform host and application integration across VMware Windows Linux/UNIX databases containers and other supported platforms.
- Analyze performance using IOPS throughput latency queue depth cache port utilization and workload patterns; remediate bottlenecks across the end-to-end I/O path.
- Execute capacity planning thin-provisioning governance reclamation tiering replication snapshot management migration refresh and decommission activities.
- Lead a meaningful engineering initiative such as automation upgrade migration recovery validation capacity optimization or security remediation with measurable outcomes.
Backup Recovery and Data Protection Engineering
- Engineer and support Cohesity Commvault Hitachi Data Protection Suite FortKnox or equivalent cyber-vault capabilities tape libraries and approved cloud data-protection services.
- Design and maintain protection policies for virtual machines physical servers databases NAS cloud workloads SaaS services and critical application data.
- Configure and optimize full incremental differential synthetic snapshot replication auxiliary copy disk-to-disk disk-to-tape and long-term retention workflows as applicable.
- Lead complex restores including granular point-in-time application-consistent bare-metal alternate-location mass recovery and disaster-recovery scenarios.
- Improve protection coverage backup success recovery confidence media management retention compliance deduplication efficiency throughput and license utilization.
First 1218 Months Success Expectations
- Stabilize and improve daily operational health across storage SAN backup and recovery services by reducing recurring incidents improving alert quality and strengthening runbook-based support.
- Deliver few measurable engineering improvements such as automation upgrade migration recovery validation capacity optimization cyber-resilience enhancement or license/cost optimization.
- Improve recoverability by validating critical restores documenting recovery evidence closing protection gaps and aligning backup policies with agreed RPO RTO retention and compliance expectations.
- Strengthening platform governance by improving documentation inventory accuracy capacity forecasting change quality lifecycle planning and vendor-support engagement.
- Contribute to team capability by mentoring engineers sharing technical knowledge participating in design reviews and promoting consistent engineering standards across the Storage & Backup capability.
Qualifications :
Education
- Bachelors degree in computer science Information Systems Engineering or a related discipline or an equivalent combination of education and relevant professional experience.
Required Experience
- 12 years of relevant infrastructure engineering experience including substantial hands-on responsibility for enterprise storage SAN backup recovery or data-resilience services. Equivalent depth and demonstrated capability may be considered.
- Experience designing implementing upgrading troubleshooting and supporting production infrastructure in a large complex or globally distributed environment.
- Hands-on experience in Pure Hitachi storage platform and Cohesity Commvault backup/data-protection platform with the ability to work across the combined capability.
- Experience contributing to technical designs and architecture decisions executing production changes resolving high-severity incidents and completing root-cause analysis.
- Practical scripting or automation experience using PowerShell Python REST APIs Ansible/AWX or comparable tools.
- Experience with Git-based workflows Agile delivery practices IT service management change control monitoring testing and modern engineering ways of working.
- Experience collaborating across teams and communicating with stakeholders vendors senior engineers architects and management.
- Exposure to modern infrastructure practices include automation AI-assisted operations observability cyber resilience cloud integration self-service and engineering governance.
Knowledge Skills and Abilities
- Strong understanding of SAN/NAS/block/file/object storage concepts multipathing zoning masking replication snapshots high availability performance and capacity management.
- Strong understanding of backup architecture retention deduplication encryption recovery methods RPO/RTO disaster recovery cyber recovery immutability isolation and recovery testing.
- Working knowledge of VMware Windows Linux/UNIX networking fundamentals DNS/NTP authentication certificates databases cloud containers and application dependencies.
- Ability to analyze complex technical problems systematically use data to validate hypotheses assess risk and recommend supportable solutions.
- Ability to balance operational work engineering delivery technical debt incidents and planned commitments while maintaining quality and control.
- Strong written verbal documentation presentation and interpersonal communication skills; ability to influence without direct authority.
- Ability to apply modern engineering practices such as infrastructure as code policy-driven automation secure-by-design delivery telemetry-based decisions and continuous improvement.
Preferred Qualifications
- Deep expertise in one or more of: Cohesity Commvault Hitachi Data Protection Suite Dell PowerStore Hitachi VSP/HNAS Pure Storage Brocade Fibre Channel SANnav tape cyber vault Azure Backup AWS Backup or comparable enterprise technologies.
- Experience with hybrid-cloud data protection Kubernetes/container protection SaaS protection object storage cloud tiering and cloud cost optimization.
- Experience with observability platforms ServiceNow Power BI capacity forecasting event integration CMDB/discovery or enterprise reporting.
- Experience with self-service infrastructure policy as code automated compliance evidence recovery orchestration AI-assisted operations or autonomous remediation.
Working Relationships
Internal Relationships
- Storage & Backup engineers; Senior/Principal Engineers Architects Engineering Managers Product and Program partners.
- Security Network Compute Cloud Database Application Platform Engineering Data Center Service Management Cybersecurity and Risk Management teams.
- Business stakeholders and internal customers whose services depend on protected and available data.
External Relationships
- Technology vendors support organizations managed service providers implementation partners and consultants.
- Industry communities and technical partners as applicable.
Working Conditions and Physical Requirements
- Primarily performed in a professional office or approved remote-work environment using standard business and collaboration technology.
- Requires collaboration across departments locations and time zones with schedule flexibility for planned maintenance technology implementations major incidents and globally distributed teams.
- May require participation in an on-call rotation and after-hours support for critical incidents or planned production activity.
- May require visits to data centers or operational facilities where approved safety and access procedures must be followed.
- Ability to perform the essential functions of the position with or without reasonable accommodation.
Additional Information :
Expeditors offers excellent benefits:
- Paid Vacation Holiday Sick Time
- Health Plan: Medical
- Life Insurance
- Employee Stock Purchase Plan
- Training and Personnel Development Program
- Growth opportunities within the company
- Employee Referral Program Bonus
Remote Work :
No
Employment Type :
Full-time
About Company
Expeditors is a Fortune 500 service-based logistics company with headquarters in Seattle, Washington, USA. At Expeditors, we generate highly optimized and customized supply chain solutions for our clients with unified technology systems integrated through a global network of over 350 ... View more