Start free Get a demo
← Careers OPEN ROLE / staff_infrastructure_engineer_device_cloud
Revyl · San Francisco

Staff Infrastructure Engineer Device Cloud

Own the technical strategy and systems required to scale the device cloud and distributed compute platform behind Revyl’s mobile development platform.

Apply to this role Read role details
Role specification 01 / DETAILS
01 about_the_role

About the role

Revyl gives developers and AI agents cloud iOS and Android environments where they build, run, inspect, and verify mobile applications. Behind that experience is a distributed compute platform spanning GCP, physical Mac infrastructure, simulators, emulators, build runners, streaming, and test execution.

We’re entering a period where our infrastructure needs to support significantly more customers, workloads, and compute. You’ll be responsible not only for operating that infrastructure, but for figuring out how it evolves: identifying bottlenecks and scaling limits, and owning the technical roadmap while moving fluidly between understanding current systems, designing what comes next, and building it hands-on.

02 what_youd_work_on

What you’d work on

  • Define the infrastructure scaling strategy: analyze the current architecture, identify what breaks under increased load, and build the roadmap to scale ahead of demand.
  • Capacity planning and modeling that translates customer and partnership growth into compute, storage, network, device, and hardware requirements.
  • Systems for provisioning, scheduling, managing, and recycling large fleets of iOS simulators, Android emulators, and physical compute.
  • Distributed workload orchestration for builds, tests, agent sessions, and interactive workloads across heterogeneous compute pools.
  • The GCP infrastructure behind our APIs, control plane, workflow execution, storage, networking, and supporting services.
  • Finding and eliminating bottlenecks with production data, from CPU, memory, disk, and network through scheduler throughput, database performance, and host capacity.
  • Preparing the platform for step-function growth with load tests, failure tests, and capacity plans for large new customers and partnerships.
03 what_were_looking_for

What we’re looking for

  • Experience taking an infrastructure platform from one stage of scale to the next.
  • 5+ years of software engineering, infrastructure, platform engineering, or SRE experience with significant technical ownership.
  • You’ve helped scale production infrastructure or a compute platform through a meaningful increase in load.
  • You can analyze an existing architecture, identify scaling constraints, and prioritize a technical roadmap.
  • You understand capacity planning and how to translate business forecasts into infrastructure requirements.
  • You make architectural decisions under uncertainty and validate assumptions with instrumentation, benchmarking, load testing, and production data.
  • You’ve built or operated distributed systems at meaningful scale.
  • You’re a strong software engineer who builds infrastructure services and control-plane systems, not just configures tooling.
  • You understand scheduling, queues, worker pools, concurrency, backpressure, resource allocation, and failure recovery.
  • Deep experience with cloud infrastructure on GCP, AWS, or Azure.
04 what_success_looks_like_in_the_first_year

What success looks like in the first year

  • We have a clear model of the capacity and architecture needed to support the next stages of growth.
  • Major scaling limits are known before a customer hits them.
  • Concrete capacity and load plans exist for large new customers and partnerships before launch.
  • The platform handles an order of magnitude more concurrent compute without a proportional increase in operational work.
  • Common infrastructure failures are detected and remediated automatically.
  • Device scheduling and capacity management are predictable instead of reactive.
  • Infrastructure projects are prioritized by real scaling constraints rather than the most recent incident.
  • The team spends substantially less time firefighting, and the Device Cloud becomes a platform that scales repeatably instead of being reinvented with each growth phase.
About Revyl 02 / COMPANY
Mobile, understood at runtime.

We are building the mobile development platform.

Revyl is the mobile development platform where teams and AI agents build, develop, test, and understand mobile applications against the real, running app.

Cloud compute and devices power the work; replayable evidence proves the result; Atlas turns every run into durable application context.

The founding team created DragonCrawl at Uber AI and left to commercialize their work. Revyl is backed by Felicis, General Catalyst, and Y Combinator, with support from the co-creator of OpenTelemetry and operators from Meta, Nvidia, and Uber.

READY / APPLY

Interested in the Staff Infrastructure Engineer — Device Cloud role?

Apply on Y Combinator with a short note and links to work you’re proud of. We read every application.

Start your application