Senior GCP Cloud / Platform Operations Engineer
Бюджет: $20.0 - $45.0
HOURLY / FULL_TIME
⭐ 4.93 (28)
United States
cloud-architecture, google-cloud-platform, python, devops, docker, terraform, cicd, infrastructure-as-code, hipaa, automated-deployment
Preferred qualifications
- Talent type: Independent
- Experience: Expert
- English: Native
🛠️ Senior GCP Cloud / Platform Operations Engineer
We're a GCP-based healthcare AI company building production automations and multi-practice cloud infrastructure. Looking for a senior cloud / platform ops engineer to keep production reliable, harden our shared GCP platform with IaC, and take day-to-day operational load off our Cloud Lead.
Hands-on build-and-run role — not a product feature developer. You partner with the Cloud Lead; you own reliability, platform hygiene, and engineer enablement so app teams can ship safely.
──────────────────────────────────────────────
🧰 Tech stack
• Compute / serverless: Cloud Run (services + jobs), Cloud Functions, Compute Engine + Docker where needed
• Eventing / data: Cloud Scheduler, Pub/Sub, Firestore, Cloud SQL (MySQL)
• Security / access: IAM, Secret Manager, IAP, VPC / private networking / egress
• Delivery: Terraform (preferred), Cloud Build and/or GitHub Actions + Workload Identity, Artifact Registry
• Ops: Cloud Monitoring / Logging, dashboards, alerts, SLOs, billing / cost hygiene
We are all-in on GCP — real Google Cloud experience is important.
──────────────────────────────────────────────
🧱 What you'll do
Keep our GCP estate reliable — Cloud Run, Functions, Scheduler, Pub/Sub, Firestore, IAM, networking — with dashboards, alerts, SLOs, incident response, and post-mortems that become durable fixes
Own Infrastructure as Code and CI/CD — Terraform for new and existing projects; repeatable deploys via Cloud Build / GitHub Actions; no snowflake console as long-term source of truth
Help stand up and operate environments for engineering and practices — IAP-protected admin surfaces, Cloud SQL, private networking, secrets and service-account design (execution under Cloud Lead)
Enable engineers safely — group-based IAM, Artifact Registry readers, shared SQL access, deploy permissions, and documented runbooks instead of one-off personal grants
Maintain cost and compliance hygiene — billing visibility, quotas, org-policy awareness, and keeping regulated workloads on a compliant path
Grow into either a practice-hosting support lane or an automations / standards lane as we hire and split work — this role is the shared GCP core both need
──────────────────────────────────────────────
✅ Must-have
• Hands-on GCP: Cloud Run, Cloud Functions, Cloud Scheduler, Pub/Sub, IAM, VPC networking, Secret Manager, Cloud Monitoring / Logging
• Infrastructure as Code — Terraform strongly preferred (Pulumi OK)
• CI/CD and deployment automation on GCP
• Incident response — diagnose and own production failures under pressure
• GCP security basics — least-privilege IAM, service accounts, secrets, private networking
• HIPAA compliance experience and/or formal training
• Clear written runbooks; comfortable pairing with a Cloud Lead
• Senior-level experience (~5+ years cloud infrastructure / DevOps / SRE); can propose architecture and work independently
➕ Nice-to-have
• Cloud SQL (MySQL) ops — HA, backups, Auth Proxy / private IP
• Identity-Aware Proxy (IAP) for admin UIs and SSH
• Firestore (incl. multi-database / app patterns)
• Docker + digest-pinned container deploys
• Cost optimization / billing export / BigQuery
• GCP Professional Cloud Architect or Cloud DevOps Engineer cert
• Open-source EHR ops experience (e.g. OpenEMR) — a plus, not a filter; we train the product surface
• Healthcare automation / RCM exposure; familiarity with common practice systems a plus
• Playwright / headed-browser infra for login automation
• Vertex AI / model-serving ops (support, not research)
──────────────────────────────────────────────
⏱️ Engagement
Full-time · remote · long-term · partner daily with our Cloud Lead Priority is high — production automations, upcoming cloud migration work, and multi-practice infrastructure all need dedicated ops capacity. Most of the team works on PST and EST time zones.
──────────────────────────────────────────────
❓ Screening questions
Describe a production incident you owned end-to-end on GCP (or similar). What failed, how you diagnosed it, and what durable fix you left behind?
Which GCP services have you operated in production day-to-day? Call out Cloud Run, IAM, networking, and monitoring specifically if applicable.
How do you keep infrastructure as code (Terraform or similar) as the source of truth when a team is moving fast?
Briefly describe your HIPAA (or equivalent regulated-data) experience in a cloud environment — training, BAAs, IAM / audit / boundary practices you've applied.
Открыть заказ
AI-черновик отклика
Короткий текст отклика для копирования в оффер: интерес + готовность работать.
Войдите, чтобы сгенерировать AI-черновик.
Войти