Site Reliability Engineer
Software Engineering
New York, NY, USA
Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.
Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.
This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.
Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.
We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.
You'll win here if…- You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio
- Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system
- You’re comfortable moving quickly, shipping improvements, and iterating in production
- AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster
- You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)
- You have experience building or orchestrating AI/agent workflows - in production or through serious side projects
Reliability & Infrastructure
- Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel
- Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards
- Manage IAM, service accounts, and security best practices across our cloud environment
- Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements
Core Infrastructure & State
- Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows
- Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments
- Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution
- Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)
- Automate infrastructure provisioning, deployments, and operational workflows
AI & Next-Gen Tooling: Agent Ops
- Build agent operations infrastructure that enables AI agents to run safely and reliably in production
- Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry
- Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop
- Own visibility into AI usage, reliability, and spend as our agent footprint scales
Cross-functional Impact
- Partner closely with engineering and product teams to maintain reliability without slowing development velocity
- Act as a force multiplier across the team — helping engineers ship faster and more safely
- 4+ years in SRE, DevOps, or infrastructure/platform engineering
- Strong, hands-on experience with a major cloud platform (preferable GCP)
- Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)
- Solid understanding of IAM, security, and cloud best practices
- Experience with observability tools like Datadog
- Familiarity with Node.js environments
- AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.
- Experience building or orchestrating AI/agent workflows (work or serious personal projects)
- High ownership, strong curiosity, and a bias toward action
- Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls
- Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus
- Hands-on experience with GCP Workflows for orchestration
- Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling
- Experience with Vercel deployment and edge computing
- Familiarity with the k8s ecosystem
- Familiarity with Redis and BullMQ
- Understanding of SOC 2 compliance requirements and implementation
- Previous experience in a high-growth startup environment
- Previous backend engineering experience to bridge the gap between infrastructure and code
- Cloud: Google Cloud Platform (GCP)
- Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes
- Database: PostgreSQL with pgbouncer, Prisma
- Observability: Datadog
- Runtime: Node.js, TypeScript
- Deployment: Vercel, GCP
- Frontend: React, Next.js, TypeScript
- Backend: TypeScript, PostgreSQL, Next.js

