Founding Senior Infra Engineer, Distributed Infra
ConfidentialAbout us We’re building toward a world where every company can become its own AI lab. Goaly is a stealth AI startup founded by ex-Meta Superintelligence Labs engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last. Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one. About the role You will own systems across the lifecycle of our accelerator clusters—from bringing capacity online and upgrading fleets to detecting failures, recovering safely, and retiring capacity. Your work will determine how quickly researchers can start experiments, how efficiently expensive hardware is used, and how reliably long-running workloads complete. This is a hands-on infrastructure role spanning cloud and datacenter environments, cluster control planes, networking, storage, security, observability, and automation. You will partner with hardware and cloud providers and our Training, RL Systems, Post-Training, Inference, and Security teams to turn heterogeneous compute into a dependable platform. Depending on experience, you may lead multi-quarter initiatives and help set technical direction. What you'll do Design, build, and operate control-plane services and infrastructure-as-code for provisioning, configuration, validation, upgrades, expansion, draining, recovery, and decommissioning; make every change repeatable, auditable, and safe to roll back. Bring new accelerator capacity online on schedule by coordinating dependencies across cloud providers, datacenter and hardware partners, networking, storage, security, and internal compute consumers. Build high-bandwidth, topology-aware connectivity within and across clusters; diagnose performance and reliability issues spanning hosts, switches, routing, transport, collective communication, and workload placement. Make clusters secure by default through identity and access controls, network policy, workload isolation, host and container hardening, secrets management, and trusted software and image supply chains. Improve fleet scalability, consistency, and fault tolerance by defining health signals, automating remediation, reducing configuration drift, and designing for partial failure. Establish service-level objectives and observability for cluster readiness, provisioning time, usable capacity, job-start latency, infrastructure-caused failures, utilization, and recovery time. Lead incident response and blameless postmortems for cluster failures; turn recurring operational pain into automation, safer defaults, and simpler system boundaries. Work directly with Training, RL Systems, Post-Training, and Inference engineers to debug cross-layer failures and shape a long-term compute, data, networking, and capacity roadmap. You may be a good fit if you have Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure). Strong programming ability in Python, Go, Rust, or another language suited to reliable infrastructure services and automation. Hands-on experience with Linux, containers, Kubernetes or another cluster scheduler, infrastructure-as-code, and at least one major cloud platform or substantial bare-metal environment. A practical understanding of networking, storage, identity, observability, and reliability, with the ability to trace a failure across multiple layers of a complex system. Experience designing systems for safe rollout, fault isolation, idempotency, capacity growth, and recovery from partial or large-scale failures. High ownership and clear communication, including comfort coordinating multi-team projects and participating in a healthy on-call rotation. Strong pluses Experience operating large GPU or acce…