Site Reliability Engineer

Aisle
Aisle

Software Engineering

New York, NY, USA

Posted on Aug 22, 2026
What is Aisle?

Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.

Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.

This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.

Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.

We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.

You'll win here if…
  1. You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio
  2. Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system
  3. You’re comfortable moving quickly, shipping improvements, and iterating in production
  4. AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster
  5. You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)
  6. You have experience building or orchestrating AI/agent workflows - in production or through serious side projects
About your role

Reliability & Infrastructure

  • Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel
  • Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards
  • Manage IAM, service accounts, and security best practices across our cloud environment
  • Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements

Core Infrastructure & State

  • Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows
  • Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments
  • Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution
  • Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)
  • Automate infrastructure provisioning, deployments, and operational workflows

AI & Next-Gen Tooling: Agent Ops

  • Build agent operations infrastructure that enables AI agents to run safely and reliably in production
  • Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry
  • Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop
  • Own visibility into AI usage, reliability, and spend as our agent footprint scales

Cross-functional Impact

  • Partner closely with engineering and product teams to maintain reliability without slowing development velocity
  • Act as a force multiplier across the team — helping engineers ship faster and more safely
About your skillsMust haves
  • 4+ years in SRE, DevOps, or infrastructure/platform engineering
  • Strong, hands-on experience with a major cloud platform (preferable GCP)
  • Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)
  • Solid understanding of IAM, security, and cloud best practices
  • Experience with observability tools like Datadog
  • Familiarity with Node.js environments
  • AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.
  • Experience building or orchestrating AI/agent workflows (work or serious personal projects)
  • High ownership, strong curiosity, and a bias toward action
Nice to haves
  • Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls
  • Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus
  • Hands-on experience with GCP Workflows for orchestration
  • Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling
  • Experience with Vercel deployment and edge computing
  • Familiarity with the k8s ecosystem
  • Familiarity with Redis and BullMQ
  • Understanding of SOC 2 compliance requirements and implementation
  • Previous experience in a high-growth startup environment
  • Previous backend engineering experience to bridge the gap between infrastructure and code
About the stack
  • Cloud: Google Cloud Platform (GCP)
  • Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes
  • Database: PostgreSQL with pgbouncer, Prisma
  • Observability: Datadog
  • Runtime: Node.js, TypeScript
  • Deployment: Vercel, GCP
  • Frontend: React, Next.js, TypeScript
  • Backend: TypeScript, PostgreSQL, Next.js