Skip to content
Back to jobs

Observability Engineer (Prometheus / Grafana / Datadog)

frontendnode-production.up.railway.app
Anywhere
Posted 1d ago
remotefull-time

Key details

Work type
remote
Employment
full time

Job Description

Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.

As we continue to grow, we’re looking for a skilled Observability Engineer (Prometheus / Grafana / Datadog) to join our dynamic team and contribute to our mission of transforming business processes through technology.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Observability Engineer (Prometheus / Grafana / Datadog) Job Title: Observability Engineer (Prometheus / Grafana / Datadog) Location: 100% Remote (Continental United States) Position Type: In-house Bright Vision Technologies SOW engagement (no third-party client or vendor) Experience: 5+ years Salary Range : $100k to $150k per annnum Sponsorship: No new H1B sponsorship available.

H1B transfers welcomed for qualified candidates.

Employment Type: Full-time, direct W2 with Bright Vision Technologies (no C2C, no 1099, no third-party) Engagement: Long-term, multi-year, aligned to the Bright Vision SOW delivery roadmap Compensation: Competitive base salary commensurate with experience, plus benefits.

Employment Terms & Visa Policy This is a 100% remote, full-time, direct W2 position with Bright Vision Technologies. This role is part of Bright Vision Technologies’ in-house Statement of Work (SOW) engagement.

The client, end customer, and employer for this position is Bright Vision Technologies — there is no third-party client, vendor, or implementation partner involved. We do not engage in C2C, 1099, or third-party arrangements for this role.

BUT STRICTLY NO C2C/1099/3RD PARTY COMPANIES. ALL OUR ROLES ARE W2 AND NO 3RD PARTY BROKERING PLEASE. Candidates must be willing to work directly as a full-time W2 employee of Bright Vision Technologies and contribute to our in-house SOW deliverables.

No new H1B sponsorship is available for this role. However, candidates who are currently on a valid H1B visa and require a transfer are welcome to apply. We will support H1B transfers for qualified candidates. For every role, a technical coding assessment is mandatory.

Please apply only if you are confident in your technical abilities and hands-on experience.

Job Summary We are looking for an Observability Engineer to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run.

The role spans the full observability stack — from collection agents and pipelines to long-term storage, dashboards, and alerting workflows — with a strong focus on usability, signal quality, and operational ROI.

The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.

Key ResponsibilitiesDesign and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring. Architect Prometheus / Thanos / Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments for high availability and scale.

Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging conventions. Define and enforce SLOs, SLIs, and error budgets, and build the dashboards and alerts that operationalize them.

Build alerting strategies that minimize noise, surface actionable signals, and integrate cleanly with on-call workflows in PagerDuty, Opsgenie, or similar tools. Operate large-scale time-series and log storage platforms, balancing retention, query performance, and cost.

Design distributed tracing pipelines and help teams use traces to diagnose latency and reliability issues. Develop self-service tooling, paved-road libraries, and templates that make adoption of observability standards easy for product teams.

Drive cost management and label-cardinality discipline across the observability estate. Lead incident response readiness improvements through better dashboards, alerting hygiene, and post-incident analysis tooling.

Partner with SRE and platform teams to integrate observability into deployment pipelines, canary analysis, and progressive delivery workflows. Evaluate and recommend observability vendors and open-source tools based on cost, capability, and operational maturity.

Mentor engineering teams on observability fundamentals, debugging techniques, and SLO-driven operations. Maintain documentation, onboarding guides, and runbooks for the observability platform. Required QualificationsBachelor’s degree in Computer Science or a related field.

Five or more years of experience in SRE, platform engineering, or observability roles. Deep hands-on experience with Prometheus, Graf

Company & context

Evidence is labeled so you can tell internal community data from public sources.

Context may refresh in the background.

Source: Google Jobs • Last updated 8h ago