Production Support & Incident Response Lead
Makati City - Philippines
Job Summary
Production Support & Incident Response Lead
When something goes wrong in production overnight you are the person leading the response.
EviSmart is looking for a Global Continuity Lead to own incident response and platform continuity during our night operations.
This is not a role where you simply monitor dashboards create a ticket and wait for Engineering.
You will be the lead incident response person on shift. You are expected to investigate first understand what is happening determine the safest way to restore operations bring in the right technical people when necessary and remain accountable for the incident until the platform is stable.
Our platform supports 2000 dental labs so an issue in production can quickly become a real business problem for our customers. The goal is simple: keep cases moving and minimize disruption.
What youll own:
You will be responsible for the health and continuity of the platform during your coverage.
That means:
Lead the response to production incidents during the night shift
Continuously monitor platform health and act on warning signs before they become customer-impacting problems
Personally perform first-line investigation using logs dashboards monitoring tools diagnostic commands and available system access
Determine the impact and likely source of an issue before escalating
Look for safe workarounds or restoration options when a permanent fix is not immediately available
Decide when an issue can be handled at your level and when Engineering DevOps or another specialist needs to be brought in
Command the incident even after technical teams become involved: keep people aligned decisions moving and communication clear
Keep Application Support and other stakeholders informed during active incidents
Document incidents root causes workarounds and follow-up actions
Make sure recurring issues dont simply become accepted problems
Provide a complete handoff to the daytime team with nothing dropped overnight
Develop one Application Support teammate into a reliable backup who can eventually handle routine night triage independently
The roles ownership of monitoring incident command proactive customer communication post-mortems and backup development is explicit in the operating playbook.
What this role is NOT
This is not a traditional Service Delivery Manager or ITIL governance position.
It is also not a pure DevOps or Software Engineering role.
You dont need to be the person who writes the permanent code fix for every problem. But you do need enough technical depth to investigate intelligently before asking someone else to solve it.
If your normal incident process is: Alert Create ticket Escalate Wait
this probably isnt the right role.
Were looking for someone whose instinct is closer to:
Detect Investigate Isolate Restore or Work Around Escalate Intelligently Command Through Resolution Prevent Recurrence
The kind of person were looking for
You may currently be an:
Senior Application Support Engineer L2/L3 Application Support Engineer Production Support Engineer Application Operations Engineer Technical Operations Engineer Platform Support Engineer or similar.
More important than your current title is how you operate.
You are someone who:
Has personally supported live production applications
Can investigate an unfamiliar production problem without immediately needing someone to tell you what to check
Is comfortable working with logs dashboards APIs databases and monitoring/observability tools
Understands enough infrastructure and application behavior to distinguish between likely application database API/connectivity and infrastructure problems
Thinks about business continuity not only technical resolution
Can make sensible decisions with incomplete information
Knows when a workaround is safer and faster than waiting for the perfect fix
Stays calm when customers are affected and several teams are involved
Communicates clearly during incidents without creating noise
Can challenge or direct technical teams when an incident needs movement
Notices patterns and asks why the same problem keeps happening
Doesnt need constant hand-holding
Is comfortable being accountable when they are the most senior incident-response person available
Heres a good way to know whether youll enjoy this role.
Its 2 AM. A production issue is preventing a customer from processing cases. The permanent fix requires an engineer who isnt immediately available.
What do you do
Were looking for someone who doesnt stop at Ill escalate it.
We want someone who starts asking:
Whats actually broken
Whats the business impact
What changed
What can I verify myself
Can I safely restore the previous working state
Is there another way to keep the customers operation moving
Who genuinely needs to be involved
What can we do now instead of waiting until morning
Thats the mindset were hiring for.
What success looks like
Your first 90 days are designed to progressively prove that we can trust you with the night.
First 30 days: Learn the platform and prove you can troubleshoot real issues and identify the correct workaround without being walked through every step.
By 60 days: Independently monitor the platform recognize warning signals and own selected production tickets through resolution.
By 90 days: Independently command night incidents from detection through restoration handle more complex issues proactively identify problems and effectively delegate to your trained backup.
Ultimately success means >99% platform health during night operations incidents declared quickly complete post-mortems no dropped night-to-day handoffs and a backup capable of independently handling night triage.
Qualifications :
Ideally you have:
Strong hands-on experience in Application Support Production Support SaaS Operations or a similar production environment
Experience supporting systems in a 24/7 or on-call environment
Real production incident troubleshooting experience
Experience with monitoring and observability tools such as Grafana Datadog CloudWatch Splunk Kibana Azure Monitor or similar
Working knowledge of logs APIs SQL/databases cloud environments and basic diagnostic tools
Experience with incident response root-cause analysis and production releases
Strong judgment around escalation risk and business continuity
The ability to communicate confidently with both technical teams and business stakeholders
Deep DevOps expertise is not required. What matters is that you can investigate intelligently understand what youre seeing take the safest action available at your level and know when specialist intervention is genuinely necessary.
Important before you apply
This is a permanent night-shift role.
This is also an Individual Contributor role not a traditional people-management position. You will act as the functional point person during night coverage and will help develop a designated backup but you will not be joining to manage a large team.
The responsibility is significant because you are the person we need to trust when the daytime team isnt around.
If youre already strong in Application or Production Support and youre looking for an opportunity where youre given more ownership more decision-making authority and the chance to become the person trusted to lead production incidents wed like to hear from you.
Apply and tell us about the toughest production problem youve personally solved.
Additional Information :
Work setup: FULL ONSITE Monday to Friday
Location: One World Place BGC Taguig City
Employment Type: Full-time Permanent
Get to know us more:
EviSmart
Work :
No
Employment Type :
Full-time
About Company
Nimbyx isn’t just another venture capital firm—we’re a launchpad for game-changers. Based in BGC, Philippines, with offices in Vancouver and Seoul, we invest in disruptive healthcare and technology companies with one bold mission: to change healthcare for good.We thrive on innovation, ... View more