Enter a job title or keyword

Senior Agentic AI Scientist

Caseware


Job Location:

Bogotá - Colombia

Monthly Salary: Not provided by the employer
Posted: 29 August 2026 (6 days ago)
Application Deadline: 2 December 2026
Vacancies: 1 Vacancy

Job Summary

Caseware is one of Canadas original Fintech companies having led the global audit and accounting software industry for over 30 years with more than 500000 users across 130 countries and available in 16 different languages. While you might not have heard of us (yet) over 36000 accounting and audit professionals list Caseware as a skill on their LinkedIn profiles!

We are building the agentic platform that powers Casewares next generation of audit and accounting products and applied science is how we measure and improve quality across our agents. This is a role for someone who thinks in experiments has built evaluations for LLM-based systems and wants to build the tools that other teams and our customers rely on to ship their own agents.

You will design and run experiments that turn ambiguous quality questions into measurable results and own the evaluation methodology that internal teams and customers depend on. You will build the AI-based capabilities behind our end-to-end agent builder including the synthetic data generation and eval builders that let teams and customers create evaluate and deploy their own agents. You will also work on something genuinely novel: self-learning self-improving systems that get better both offline and online. A significant part of that is agentic memory: deciding what is worth remembering and validating the criteria that let agents compound useful knowledge on behalf of customers over time while rigorously upholding the legal and contractual obligations owed to them and their clients. At the staff level you will also influence technical direction and mentor others.

Location: This is a fully remote position located in Colombia.

Contact

Maira Russo- Senior Talent Acquisition Partner

The domain: financial audit

This role sits squarely in the financial audit domain and that domain matters more than most people realize. Independent audit is one of the quiet foundations of the global economy. When investors lenders regulators and the public can trust that a companys financial statements are accurate capital can flow markets can function and organizations can be held accountable. That trust rests on the quality of audit and assurance work which makes the tools that audit professionals use genuinely consequential.

You do not need prior knowledge of accounting or auditing to succeed here. What you will gain is deep durable expertise in how audits are performed and where AI can make them faster more reliable and more insightful. You will learn the domain on the job alongside experienced domain experts and that fluency will make you a stronger applied scientist. If you already bring that background even better.

What youll be doing:
  • Design and run experiments that measure and improve the quality of LLM-based applications and agents turning ambiguous quality questions into measurable reproducible results.

  • Build out and own the evaluation methodology in practice that internal teams and customers depend on: benchmarks regression suites scoring methods acceptance thresholds across task success faithfulness safety latency and cost and the scientific foundations for comparing offline and online performance partnering with platform engineering on the automated infrastructure that runs those comparisons and alerts on drift or regression and wiring the signals into release gates.

  • Build AI-based tools and capabilities that power the end-to-end agent builder developer experience used by other Caseware teams and by our customers.

  • Build the synthetic data generation and eval-builder capabilities that let internal teams and customers create evaluate and deploy their own agents.

  • Turn domain procedures into task-plus-verifier structures drawing on deep intuition for how LLM-based systems fail.

  • Contribute to applied science on agentic memory: use statistical methods grounded in audit domain knowledge to surface which experiences are worth remembering and define and validate the promotion and demotion criteria that let agents compound that knowledge over time partnering with platform engineering on the infrastructure that executes it.

  • Advance self-learning self-improving systems that get better both offline and online as they are used.

  • Help ensure memory promotion and demotion criteria uphold the legal and contractual obligations owed to customers and their clients partnering with Security Legal and Domain SMEs.

  • Design experiments and evaluations that demonstrate agent quality improves the more customers use the platform.

  • Translate advances in GenAI (models retrieval agent frameworks evaluation techniques) into practical maintainable capabilities.

  • At the staff level influence technical direction through RFCs and design reviews and mentor other scientists and engineers.

What youll bring:
  • 6 years (senior) to 8 years (staff) of professional experience in applied science machine learning research or data-intensive engineering. Staff-level candidates bring demonstrated impact beyond a single team.

  • 2 years working on production GenAI or LLM-based systems (applications agents or the tooling and evaluation around them).

  • Proven experience designing and building evaluation frameworks for ML or LLM systems: metrics benchmarks scoring approaches and rigorous experiment design.

  • A strong experimentation mindset. You form clear hypotheses design sound experiments and draw defensible conclusions from noisy real-world data.

  • Ability to build and operate your own production tooling and services ideally on AWS.

  • Strong understanding of GenAI system tradeoffs including quality latency cost reliability and safety.

  • A strong foundation in probability and statistics: experimental design significance testing and reasoning under uncertainty.

  • Strong English language communication and collaboration skills.

  • Comfortable operating in fast-moving environments with ambiguity and evolving requirements.

Nice to have
  • PhD or MS in a quantitative field (Computer Science Statistics Machine Learning or similar) preferred though equivalent industry experience is welcome.

  • Prior machine learning experience (classical ML model training or MLOps).

  • Experience with AI guardrails governance or safety mechanisms.

  • Experience with distributed SaaS cloud-native multi-tenant platforms at scale.

  • Familiarity with Infrastructure as Code (CDK CloudFormation or Terraform).

  • Experience operating in regulated or compliance-heavy domains.

  • Familiarity with financial audit accounting or assurance workflows.

Tech stack youll be working with
  • Backend & Platform: TypeScript NestJS Python

  • Cloud & Infrastructure: AWS EKS AWS Lambda AWS Bedrock AWS AgentCore

  • Search & Retrieval: AWS OpenSearch and S3 Vectors

  • Document & Data Processing: AWS Textract DynamoDB S3

  • AI Evaluation & Observability: LangFuse LangSmith LangChain LangGraph

  • AI-Assisted Development: GitHub Copilot Claude Code Devin

  • Developer Tooling: GitHub GitHub Actions Nx Monorepo

Perks & Benefits
  • Contrato a termino Indefinido with all the legal benefits
  • Prepaid Medicine
  • Life insurance and funeral assistance
  • Internet allowance
  • Home office stipend
  • Competitive compensation above the market average
  • 100% remote work environment and an excellent work-life balance
  • 5 Personal Time Off days per year
  • Sick Leave Top up to total 100% of salary paid by the employer from Day 3 to 90.
  • Recognition Award additional paid time off in recognition of the corresponding year of service
  • Upgrade vacation starting at 5 years of service
  • Opportunity to work for a growing global SaaS leader company
  • A culture that promotes independence innovation trust and accountability
  • Open space to be creative innovative and strategize for the future
  • Mentorship by highly experienced professional
  • Budget for training we want you to grow
  • AI-first environment: Be part of an AI-first engineering organization that embraces modern tools automation and AI-driven ways of working.
Whats in it for you:
Innovation is at our core. We work with cutting-edge technology in accounting and financial reporting constantly pushing the boundaries to create impactful software solutions.
We are committed to a collaborative culture where your ideas are valued and knowledge sharing is encouraged within a supportive inclusive team.
Work-life balance is important to us. We offer flexible work options remote opportunities and generous time-off policies to ensure a healthy work-life balance.
We offer competitive compensation including a competitive salary and comprehensive benefits
We are driven by impactful work. Your contributions directly affect how our clients manage financial processes and drive their success.
Recognition and rewards matter to us. We celebrate hard work through recognition programs performance bonuses and opportunities for career growth.
We embrace global opportunities. Work on international projects and collaborate with a diverse global team.
About Caseware:
Casewares cutting-edge software products are meticulously designed for accounting firms corporations and teams are continually collaborating innovating and building upon our existing suite of products. With a customer-focused mindset we are building technology that is shaping what the future of audits financial reporting and financial data analytics will look like.
With a recent strategic investment from Hg Capital in 2020 Caseware is now in its next major growth phase as we double down on the people and products that have made Caseware so successful to date.
One of Casewares core values is Many Voices One Team and with that in mind were dedicated to building teams as diverse as our customers in an equitable and inclusive way. We welcome and encourage candidates of all backgrounds to apply. Should you require accommodations or have any questions at any point during the application or interview process please e-mail our People Operations team at emailprotected.
Background Check:
Any candidates successful in obtaining an offer for a position will need to successfully complete a background check through which typically includes an Identity Verification and Criminal Record Check. Executives and Senior Managers will undergo a Soft Credit Check as well. Candidates residingin the Netherlands and Germany are excluded from undergoing background checks via
Security and Fraud:
Caseware takes the security of candidates seriously. All legitimate communication from us will come from email addresses ending in @ and our open positions are always listed on reputable job boards and on our website We will NEVER ask for payment or financial information from you. If you receive an unsolicited job offer proceed with extreme caution.
We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications analyzing resumes or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed please contact us.

Required Experience:

Senior IC


About Company

Company Logo

Accounting, audit, analytics and compliance software built by seasoned accountants. Manage your audit and financial reporting more efficiently with less risk.

View Profile View Profile