Today theres more data and users outside the enterprise than inside causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed one that is built in the cloud and follows and protects data wherever it goes so we started Netskope to redefine Cloud Network and Data Security.
Since 2012 we have built the market-leading cloud security company and an award-winning culture powered by hundreds of employees spread across offices in Santa Clara St. Louis Bangalore London Paris Melbourne Taipei and Tokyo. Our core values are openness honesty and transparency and we purposely developed our open desk layouts and large meeting spaces to support and promote partnerships collaboration and teamwork. From catered lunches and office celebrations to employee recognition events and social professional groups such as the Awesome Women of Netskope (AWON) we strive to keep work fun supportive and interactive.Visit us atNetskope Careers. Please follow us on LinkedIn and Twitter@Netskope.
As a Senior AI Quality & Red Team Engineer at Netskope you will lead the charge in testing stress-testing and breaking our AI agents before they ever reach production. From automating multi-turn prompt injections to tracking fleet-wide drift in CI/CD you will own the automated harness that ensures our AI systems are secure resilient and compliant. If you love the idea of being the person who proves an agent isnt ready yet welcome home.
Build and grow the automated evaluation suite every agent runs against before its approved for production designed to run unattended and scale across a growing agent fleet not something that needs a person babysitting each run.
Design adversarial test scenarios prompt injection attempts sycophancy checks where an agent has to correctly push back on a false premise multi-attempt attacks rather than single-shot ones and automate them so they run on every relevant change not just before a big release.
Own the break it on purpose pass for every new agent: attempt to extract data it shouldnt expose get it to act outside its registered tool boundaries or get it to treat a synthetic test probe as real. As the fleet grows build this into a repeatable scriptable process rather than a manual exercise redone from scratch each time.
Partner with the Data Steward on data sensitivity classification for the systems agents touch so your test scenarios reflect whats actually at stake not a generic checklist.
Decide for each agent capability what pass actually means and build that judgment into automated thresholds wherever possible so evaluation keeps up as the number of agents climbs into the hundreds.
Maintain the guardrail and negative-test catalog (fail-closed vs. fail-graceful behavior) across the platform and add new cases as new failure modes get discovered in the wild.
Produce clear audit-ready evidence for every agents evaluation results generated automatically as part of the pipeline rather than assembled by hand for each review.
Track drift over time across the whole fleet not agent by agent so a slow-moving problem in one corner doesnt go unnoticed just because no ones looking at that specific agent that week.
Must-Have:
At least 4 years in software quality security testing or a related discipline with 12 years specifically evaluating or red-teaming LLM-based systems not just running unit tests against traditional code.
Strong Python skills since the evaluation harness adversarial test scripts and automated pipelines will mostly be built in it. Comfortable writing production-quality code not just glue scripts.
Working knowledge of REST APIs and webhook/event-driven patterns enough to build test harnesses that call an agents tools directly and validate its inputs and outputs not just its final chat response.
Real experience building automated test infrastructure and integrating it into CI/CD not just manually running test cases someone who thinks in pipelines and repeatability by default.
Hands-on familiarity with at least one adversarial testing or LLM eval tool (DeepTeam Garak PyRIT Promptfoo or similar) and an understanding of how these map to standards like the OWASP LLM Top 10 or NISTs AI risk framework.
Practical understanding of prompt injection jailbreaking and sycophancy failure modes able to design new test cases for these not just run ones someone else wrote.
Comfort making a hard call: willing to block a release when an agent doesnt meet its bar even under schedule pressure.
Strong Advantage:
Enough understanding of data sensitivity and compliance classification to design tests that reflect real risk even though the Data Steward owns the classification system itself.
Experience in a regulated environment where evaluation results had to hold up to an external audit not just an internal review.
Familiarity with multi-attempt or persistent-attack testing methodology rather than only single-shot adversarial prompts.
Some exposure to how agents are actually built (prompting tool schemas orchestration) not required but it makes it much easier to design tests that target real failure modes instead of generic ones.
#LI-CV1
Netskope is committed to implementing equal employment opportunities for all employees and applicants for employment. Netskope does not discriminate in employment opportunities or practices based on religion race color sex marital or veteran statues age national origin ancestry physical or mental disability medical condition sexual orientation gender identity/expression genetic information pregnancy (including childbirth lactation and related medical conditions) or any other characteristic protected by the laws or regulations of any jurisdiction in which we operate.
The application window for this position is expected to close within 50 days. You may apply by filling out the below information or visiting ourNetskope Careers site.
Required Experience:
Unclear Seniority
About NetskopeToday theres more data and users outside the enterprise than inside causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed one that is built in the cloud and follows and protects data wherever it goes so we started Netskope to redefine Cloud Net...
About Netskope
Today theres more data and users outside the enterprise than inside causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed one that is built in the cloud and follows and protects data wherever it goes so we started Netskope to redefine Cloud Network and Data Security.
Since 2012 we have built the market-leading cloud security company and an award-winning culture powered by hundreds of employees spread across offices in Santa Clara St. Louis Bangalore London Paris Melbourne Taipei and Tokyo. Our core values are openness honesty and transparency and we purposely developed our open desk layouts and large meeting spaces to support and promote partnerships collaboration and teamwork. From catered lunches and office celebrations to employee recognition events and social professional groups such as the Awesome Women of Netskope (AWON) we strive to keep work fun supportive and interactive.Visit us atNetskope Careers. Please follow us on LinkedIn and Twitter@Netskope.
As a Senior AI Quality & Red Team Engineer at Netskope you will lead the charge in testing stress-testing and breaking our AI agents before they ever reach production. From automating multi-turn prompt injections to tracking fleet-wide drift in CI/CD you will own the automated harness that ensures our AI systems are secure resilient and compliant. If you love the idea of being the person who proves an agent isnt ready yet welcome home.
Build and grow the automated evaluation suite every agent runs against before its approved for production designed to run unattended and scale across a growing agent fleet not something that needs a person babysitting each run.
Design adversarial test scenarios prompt injection attempts sycophancy checks where an agent has to correctly push back on a false premise multi-attempt attacks rather than single-shot ones and automate them so they run on every relevant change not just before a big release.
Own the break it on purpose pass for every new agent: attempt to extract data it shouldnt expose get it to act outside its registered tool boundaries or get it to treat a synthetic test probe as real. As the fleet grows build this into a repeatable scriptable process rather than a manual exercise redone from scratch each time.
Partner with the Data Steward on data sensitivity classification for the systems agents touch so your test scenarios reflect whats actually at stake not a generic checklist.
Decide for each agent capability what pass actually means and build that judgment into automated thresholds wherever possible so evaluation keeps up as the number of agents climbs into the hundreds.
Maintain the guardrail and negative-test catalog (fail-closed vs. fail-graceful behavior) across the platform and add new cases as new failure modes get discovered in the wild.
Produce clear audit-ready evidence for every agents evaluation results generated automatically as part of the pipeline rather than assembled by hand for each review.
Track drift over time across the whole fleet not agent by agent so a slow-moving problem in one corner doesnt go unnoticed just because no ones looking at that specific agent that week.
Must-Have:
At least 4 years in software quality security testing or a related discipline with 12 years specifically evaluating or red-teaming LLM-based systems not just running unit tests against traditional code.
Strong Python skills since the evaluation harness adversarial test scripts and automated pipelines will mostly be built in it. Comfortable writing production-quality code not just glue scripts.
Working knowledge of REST APIs and webhook/event-driven patterns enough to build test harnesses that call an agents tools directly and validate its inputs and outputs not just its final chat response.
Real experience building automated test infrastructure and integrating it into CI/CD not just manually running test cases someone who thinks in pipelines and repeatability by default.
Hands-on familiarity with at least one adversarial testing or LLM eval tool (DeepTeam Garak PyRIT Promptfoo or similar) and an understanding of how these map to standards like the OWASP LLM Top 10 or NISTs AI risk framework.
Practical understanding of prompt injection jailbreaking and sycophancy failure modes able to design new test cases for these not just run ones someone else wrote.
Comfort making a hard call: willing to block a release when an agent doesnt meet its bar even under schedule pressure.
Strong Advantage:
Enough understanding of data sensitivity and compliance classification to design tests that reflect real risk even though the Data Steward owns the classification system itself.
Experience in a regulated environment where evaluation results had to hold up to an external audit not just an internal review.
Familiarity with multi-attempt or persistent-attack testing methodology rather than only single-shot adversarial prompts.
Some exposure to how agents are actually built (prompting tool schemas orchestration) not required but it makes it much easier to design tests that target real failure modes instead of generic ones.
#LI-CV1
Netskope is committed to implementing equal employment opportunities for all employees and applicants for employment. Netskope does not discriminate in employment opportunities or practices based on religion race color sex marital or veteran statues age national origin ancestry physical or mental disability medical condition sexual orientation gender identity/expression genetic information pregnancy (including childbirth lactation and related medical conditions) or any other characteristic protected by the laws or regulations of any jurisdiction in which we operate.
The application window for this position is expected to close within 50 days. You may apply by filling out the below information or visiting ourNetskope Careers site.
Netskope, a global cybersecurity leader, is redefining cloud, data, and network security to help organizations apply zero trust principles to protect data.