ТРТ
Monitoring & Incident-Triage Engineer (Mid-Level)
- Google Cloud Platform
- Cloud Monitoring
- Logging
- Error Reporting
- Cloud Composer
- Airflow
- SQL
- Spanner
- BigQuery
- Python
- Docker
- Terraform
- monitoring-as-code
- Английский — C1 — Продвинутый
Stealth-mode AI-powered Cloud-Native Health-Tech company is looking for a diligent Mid-Level engineer to be the daily set of eyes on the operational health of platform on Google Cloud to join the team.
It’s not vaporware, their platform supports Physician Networks (IPAs) by enabling Smarter, Risk-Adjusted, and more Predictive Care that improves real patient outcomes.
This is not a support desk, and it is not a build-features role: it is the detection, triage, and escalation layer for production incidents, plus steady work through the backlog of small defects.
Compensation and Benefits:
Paid Time Off
The company has Unlimited PTOs Policy and compensated New Years Holidays on top of that. The misuse of the policy isn’t welcomed, though it’s definitely possible to take at least two weeks – and fully compensated – vacation, or more.
Corporate Hardware
The company provides Corporate Hardware for employees who completed their Probationary Period, as well as proven their value.
Team Building
The company partially compensates Team Building events, when multiple teammates are located in nearby countries.
Target Stack:
NOTE: Similar Cloud-Native Experience is always an option.
NOTE: Deep engineering ownership is NOT required for this role — operational familiarity with these tools is what matters.
-
Necessary: Google Cloud Platform; Cloud Monitoring / Logging / Error Reporting (reading & triaging); Cloud Composer (Airflow) — reading DAG runs & logs; Google Kubernetes Engine — reading pod status & logs; SQL (investigative queries in Spanner / BigQuery); Git (GitHub, GitFlow); Slack (alert channels).
-
As a plus: Python (reading code, small defect fixes); BigQuery / Cloud Spanner; Pub/Sub; Docker; Terraform / monitoring-as-code.
Requirements
Monitoring & Incident-Triage Engineer in US-centric Health-Tech
The core of the job is disciplined, repetitive vigilance: watch the dashboards and alerts every day across two production systems, catch failures early — classify them, and escalate to the right engineers through a defined incident line (L1/L2). You are not expected to root-cause and fix every deep issue yourself; you are expected to make sure nothing slips through unnoticed, that the right person is pulled in fast, and that recurring failures get driven down over time.
Alongside incident triage, you will work through the steady stream of simple, well-scoped follow-up defects that accumulate around both systems — the kind that are individually small but need someone reliable to pick them up, close them in batches, and keep the backlog from growing.
This role suits someone who is comfortable with a Google Cloud environment (Cloud Monitoring/Logging, Airflow/Composer UI, Kubernetes pod logs), can read a dashboard and a log, run a check from a runbook, write a clear defect report, and grind through follow-ups without losing attention to detail.
Incident Lines & Reaction-Time Targets
We run a two-line incident model. This role is the first line: detect, triage, and escalate. Second line (senior engineers / the relevant team) owns the deep fix. Targets below are for detection and escalation during working hours (to be finalized with the team):
-
L1 — this role: notice the alert / failed run, confirm it’s real (not a flake), classify severity, and escalate to the right owner with a clear summary. Target: acknowledge and triage a fired alert within 30 minutes during working hours; escalate anything user- or PHI-impacting immediately.
-
L2 — escalation target: senior engineers/ the owning team take the diagnosis and fix. This role tracks it to closure and confirms recovery but does not need to solve it.
-
Recurring or repeated failures are logged and raised, so they get a permanent fix or a better alert — the long-term goal is fewer incidents, not just faster handling.
Note: this role does not carry a formal 24/7 on-call pager. It is daytime monitoring, triage, and defect follow-up. If the company later adds a paid on-call rotation, that would be scoped and compensated separately.
Overview of Future Responsibilities:
Daily Monitoring
-
Watch the dashboards and alert channels for both systems every working day; treat it as a routine, not an afterthought.
-
Actively look for silent failures — a pipeline that succeeded but wrote nothing, a vendor path that quietly returned zero, a DAG that went green with tasks skipped.
-
Confirm whether a fired alert is real or a flake before acting.
Incident Triage & Escalation (L1)
-
Classify each real issue by severity and impact (user-facing? PHI-affecting? one tenant or fleet-wide?).
-
Escalate to the right owner (L2) with a clear, self-contained summary: what broke, when, blast radius, and what you already checked.
-
Track escalated incidents to closure and confirm recovery; own the incident timeline and communication in the alert channel.
-
You are the detection and coordination layer — deep diagnosis and the fix belong to L2.
Defect Follow-up
-
Work through the backlog of small, well-scoped follow-up defects across both systems; close them in batches and keep the backlog from growing.
-
Reproduce, isolate, and clearly document defects; verify fixes and confirm they don’t regress.
Continuous Improvement
-
Track recurring failures and raise them so they get a permanent fix or a better alert — drive the failure rate down over time.
-
Flag alerting blind spots (something broke with no alert) and noisy alerts (fires with no action needed).
-
Use and help keep the runbooks accurate as the systems change.
Required experience:
Monitoring & Triage (Mid-Level)
-
Mid-level experience in monitoring, QA, support-engineering, NOC, or junior-SRE-style roles for production software.
-
Comfort reading dashboards, metrics, and logs and telling a real failure from noise.
-
Clear defect/incident reporting: reproduction, isolation, impact, and a summary an engineer can act on.
-
SQL (intermediate) to investigate data during triage.
-
Basic Python (able to read code and make small, well-scoped defect fixes) — a plus, not a hard requirement.
-
Operational familiarity with a major cloud; Google Cloud (Cloud Monitoring, Logging, Composer/Airflow, GKE) preferred.
-
Diligence and consistency: this is deliberately repetitive work that rewards attention to detail day after day.
General Skills
-
Severity/impact judgment and a sense of when to escalate immediately vs. batch.
-
Defect lifecycle management: reproduction, isolation, clear reporting, regression verification.
-
Clear written communication under mild time pressure (incident summaries, stakeholder updates in Slack).
-
Comfort in a Unix-like environment (macOS/Linux).
Cloud-Native & GCP
-
Cloud-native familiarity is preferred. Candidates comfortable in Google Cloud Platform (GCP) will be prioritized, although experience with other major cloud providers (e.g., AWS or Azure) is also valuable.
-
This role requires working in a Unix-like Development Environment (e.g., macOS, Linux). We do not use Windows-based workstations for Engineering or AI-related tasks. Virtualization (e.g., WSL over Windows) isn’t enough.
Business Domain
-
Experience in Healthcare, Health-Tech, and MedTech is a significant advantage
Fundamentals
-
Ability to work in an Iterative Development workflow, delivering in small, reviewed increments rather than a Waterfall-style approach.
-
Experience collaborating with engineers, QA, and analysts to turn a requirement into working, tested, deployed code.
-
There are many Experience Advantages a candidate may have, e.g., familiarity with PagerDuty/Opsgenie-style tooling, Airflow operations, webhook-heavy and communication platforms (SMS/voice), AI/LLM-based systems, and basic scripting to speed up repetitive checks.
Компания, занимающаяся разработкой облачных технологий в сфере здравоохранения ищет для своей команды ответственного инженера, который будет ежедневно контролировать работоспособность облачной платформы в сфере здравоохранения на Google Cloud.
Это не «виртуальный продукт» — их платформа поддерживает сети врачей (IPA), обеспечивая более интеллектуальное, с учетом рисков и прогнозируемое медицинское обслуживание, которое улучшает реальные результаты лечения пациентов.
Это не служба технической поддержки и не должность, связанная с разработкой новых функций: речь идет об обнаружении, сортировке и эскалации инцидентов в производственной среде, а также постоянная работа над устранением накопившихся мелких дефектов.