···
Log in / Register

Quality Engineer - AI Agents, Evals & Integrations

Indeed
Full-time
Onsite
No experience limit
No degree limit
111411, Los Mártires, Bogotá, Colombia
Favourites
Share

Description

Summary: Seeking a Senior Quality Engineer to ensure reliability, safety, and correctness of AI-powered billing and collections workflows through automated evaluations and robust testing practices. Highlights: 1. Focus on AI agent evaluations, integrations, and backend services. 2. Define "good" for AI-enabled workflows and build automated evaluations. 3. Work with imperfect requirements and non-deterministic AI behavior. **Description:** ---------------- We are looking for a Senior Quality Engineer to help ensure the reliability, safety, and business correctness of AI\-powered billing and collections workflows. This role is focused on AI agents, evaluations, integrations, backend services, APIs, event\-driven workflows, and production observability. The person will help define what “good” looks like for AI\-enabled workflows, build automated evaluations that detect meaningful regressions, investigate failures across application behavior and cloud evidence, and create confidence as changes move quickly from refinement through production. This is not primarily a browser automation or end\-of\-cycle manual testing role. Browser\-based end\-to\-end coverage is important, but the core of this position is validating AI\-agent behavior and the financial, operational, and integration workflows around it. The ideal candidate is comfortable working with imperfect requirements, non\-deterministic AI behavior, external systems, asynchronous processing, and high\-consequence workflows where incorrect outputs can affect invoices, collections activity, customer experience, or operational decisions. **Requirements:** ----------------- **What You Will Do** * Define and maintain an evaluation strategy for AI\-powered workflows, including agent behavior, output quality, accuracy, completeness, safety, and business\-rule compliance. * Build automated AI evaluations using tools such as DeepEval, Ragas, Promptfoo, LangSmith, Langfuse, or similar frameworks. * Create and maintain representative evaluation datasets, golden examples, regression suites, synthetic test data, and edge\-case scenarios. * Combine deterministic checks, business rules, structured assertions, LLM\-as\-a\-judge techniques, human\-review workflows, and production evidence where appropriate. * Integrate evaluations into CI/CD pipelines so meaningful AI regressions are detected before release. * Establish practical quality thresholds, release signals, and investigation workflows for changes to prompts, models, retrieval, tools, agent logic, data, and integrations. * Test APIs, backend services, data flows, third\-party integrations, asynchronous workflows, retries, failure handling, and downstream system outcomes. * Validate billing and collections workflows end to end, including workflow state, data transformations, permissions, external\-system behavior, and customer\-impacting outcomes. * Use Langfuse, CloudWatch, logs, traces, dashboards, and data evidence to investigate failures and distinguish product defects from model behavior, environment issues, test defects, or accepted risk. * Partner with Product, Engineering, Operations, and AI platform teams to clarify expected behavior, identify risks early, and define meaningful validation coverage. * Build and maintain targeted automated API, integration, and end\-to\-end tests. Use Playwright where browser\-level coverage is the right way to protect an important workflow. * Turn escaped defects, operational issues, and production learnings into stronger evaluations, regression coverage, observability, or delivery guardrails. * Provide clear release\-readiness input based on evaluation results, integration evidence, known risk, production signals, and remaining uncertainty. Required Experience * 5\+ years of hands\-on Quality Engineering, SDET, Software Engineering, AI Quality, or equivalent experience. * Demonstrated experience testing systems with LLM, AI, ML, RAG, or agentic components. * Hands\-on experience designing, implementing, or maintaining automated AI evaluations. Experience with DeepEval is strongly preferred; comparable experience with Ragas, Promptfoo, LangSmith, TruLens, Arize, or similar tools is acceptable. * Strong understanding of how to evaluate non\-deterministic systems, including hallucinations, partial correctness, inconsistent outputs, tool\-use failures, retrieval quality, prompt regressions, and model changes. * Experience using AI observability or tracing tools such as Langfuse, LangSmith, Arize Phoenix, Helicone, or equivalent. * Strong experience testing APIs, backend services, integrations, and data\-driven workflows. * Experience testing asynchronous or event\-driven systems, including queues, retries, webhooks, scheduled jobs, background processing, and downstream effects. * Strong practical experience with AWS and CloudWatch, including investigating application behavior through logs, metrics, traces, alarms, and cloud\-service evidence. * Strong programming ability in Python and/or TypeScript or JavaScript. The candidate must be able to write maintainable automated evaluations, API tests, test utilities, and data\-validation scripts. * Experience with API testing and automation using tools such as pytest, Requests, HTTPX, Postman, Bruno, Supertest, Playwright API testing, or equivalent. * Experience using SQL to validate data persistence, workflow state, transformations, and downstream outcomes. * Experience integrating automated tests and evaluations into CI/CD pipelines. * Ability to independently investigate complex failures across application behavior, API payloads, cloud logs, data records, integrations, and AI traces. Required Mindset * Treat AI evaluation as an engineering discipline, not a one\-time prompt check. * Focus on customer and business impact rather than generic test counts. * Know when deterministic assertions are required and when probabilistic or judge\-based evaluation is appropriate. * Be comfortable challenging unclear requirements and turning them into practical, testable expectations. * Communicate uncertainty clearly and make release risk visible without becoming a delivery bottleneck. * Use AI\-assisted engineering tools productively while reviewing and validating their output before it affects code, tests, or release decisions. Secondary but Important Experience * Strong hands\-on experience with Playwright and TypeScript for browser\-based end\-to\-end testing. * Experience validating customer\-facing workflows involving authentication, permissions, account configuration, feature flags, and tenant boundaries. * Contract testing experience using Pact or similar tools. * Experience with performance, reliability, resilience, API security, or OWASP\-focused validation. * Experience with BrowserStack or similar cross\-browser testing platforms. * Experience testing financial, billing, invoicing, collections, operational, or other customer\-critical workflows. * Experience with feature rollouts, rollback planning, production smoke checks, phased releases, and post\-release monitoring. Primary Tooling Areas The strongest candidates will have practical depth across most of these areas: * AI evaluations: DeepEval, Ragas, Promptfoo, LangSmith, TruLens, custom evaluation frameworks * AI observability: Langfuse, LangSmith, traces, prompt/version analysis, cost and latency monitoring * Cloud and production evidence: AWS, CloudWatch, X\-Ray, logs, dashboards, alarms, metrics * Backend and integrations: REST APIs, webhooks, queues, event\-driven systems, data transformations, third\-party integrations * Automation: Python, pytest, TypeScript, API automation, CI/CD, GitHub Actions or equivalent * Data validation: SQL, structured\-data validation, synthetic test data, evaluation datasets * End\-to\-end coverage: Playwright and browser\-based workflow validation where it protects meaningful customer risk What Success Looks Like Within the first several months, this person will have: * Established a practical evaluation approach for the highest\-risk AI\-agent workflows. * Created repeatable regression suites and representative evaluation datasets. * Added AI evaluation and integration signals into the delivery pipeline. * Improved visibility into agent failures through Langfuse, CloudWatch, logs, traces, and dashboards. * Strengthened end\-to\-end confidence in billing and collections workflows across APIs, integrations, data flows, and user\-facing behavior. * Helped the team move faster by identifying risk earlier, reducing avoidable rework, and making release decisions more evidence\-based

Source:  indeed View original post
Valentina Rodríguez
Indeed · HR

Company

Indeed
Valentina Rodríguez
Indeed · HR

Similar jobs

Cookie
Cookie Settings
Our Apps
Download
Download on the
APP Store
Download
Get it on
Google Play
© 2025 Servanan International Pte. Ltd.