REMOTEFULLTIME
Analista de Observabilidade Sênior - Vaga Afirmativa para Mulheres
Jobgether
Remote · remote · Posted 1d ago
Your match
Sign in to see your match score, skill gaps & tailored resume.
Section · 01
About this role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Analista de Observabilidade Sênior - Vaga Afirmativa para Mulheres based in Brazil.
This is a senior technical opportunity to shape and lead the evolution of an organization-wide observability strategy. You will act as a key reference for Datadog, helping engineering and operations teams build more reliable, resilient, and proactive systems. The role covers metrics, logs, traces, application performance, infrastructure monitoring, user experience, and business indicators. You will transform operational data into actionable insights that improve platform stability, performance, and customer experience. Working across Engineering, Architecture, Development, SRE, and Operations, you will influence standards and promote observability best practices. You will also lead initiatives around automation, intelligent monitoring, incident prevention, and self-healing capabilities. This is a high-impact role for someone who combines deep technical expertise with strong communication, influence, and a continuous-improvement mindset.
Accountabilities:: Act as the technical and functional reference for Datadog across the organization, supporting its implementation, administration, governance, and continuous evolution.
Design, implement, and improve observability solutions using Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Real User Monitoring (RUM), Synthetic Monitoring, and Continuous Testing.
Define and maintain enterprise observability standards for applications, APIs, microservices, and cloud workloads.
Build and maintain executive, operational, and analytical dashboards to provide visibility into platform health and performance.
Develop proactive monitoring strategies based on SLIs, SLOs, SLAs, and relevant business indicators.
Create, review, and optimize intelligent monitors and alerts to reduce operational noise, false positives, and unnecessary escalations.
Support development teams with application instrumentation using OpenTelemetry and native Datadog integrations.
Lead root-cause analysis (RCA) for critical incidents and recommend structural improvements to prevent recurrence.
Identify opportunities for automation, predictive failure detection, automated incident response, and self-healing.
Conduct periodic assessments of capacity, performance, availability, reliability, and end-user experience.
Develop training materials, playbooks, standards, and technical documentation related to observability and monitoring practices.
Promote a culture of observability, reliability, and operational excellence across technical teams.
Translate technical indicators and operational events into business impact and actionable recommendations.
Influence engineering practices and reliability standards across multiple teams and contribute to the organization's broader observability maturity.
Requirements:
Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
Advanced hands-on experience with Datadog, including implementation, administration, configuration, and platform evolution.
Strong practical experience with Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Monitors, and Service Catalog.
Advanced understanding of the core observability pillars: logs, metrics, and distributed tracing.
Experience monitoring APIs, microservices, and distributed architectures.
Professional experience working with AWS and cloud environments.
Strong troubleshooting and complex incident investigation capabilities.
Experience with OpenTelemetry and application instrumentation.
Knowledge of automation using Python, Shell Script, or PowerShell.
Experience with Kubernetes, Docker, and cloud-native ecosystems.
Solid understanding of system availability, performance, scalability, and reliability.
Ability to connect technical metrics and operational events with business outcomes.
Excellent communication skills and the ability to collaborate with multiple technical and business stakeholders.
Investigative, analytical, and problem-solving mindset with a strong focus on continuous improvement.
Strong sense of ownership and operational responsibility.
Ability to influence technical teams and encourage the adoption of observability and reliability best practices.
Collaborative and consultative approach, with the ability to train and enable other teams.
Experience with SRE practices is desirable.
Datadog Certified Associate certification or higher is a plus.
Experience designing observability strategies for large-scale distributed environments is an advantage.
Knowledge of CI/CD and DevSecOps practices is desirable.
Experience with automated incident response and self-healing processes is a plus.
Familiarity with ITIL, incident management, Problem Management, and operational governance is desirable.
Experience with tools such as Dynatrace, Grafana, Prometheus, Elastic Stack, or Zabbix is an advantage.
Benefits:
Senior-level opportunity with significant influence over enterprise observability and reliability strategy.
High-visibility work impacting multiple engineering and technology teams.
Opportunity to lead automation, intelligent observability, and incident-reduction initiatives.
Professional development through exposure to advanced cloud-native technologies and modern observability practices.
Opportunity to work with Datadog, OpenTelemetry, AWS, Kubernetes, Docker, and related technologies.
Collaborative environment involving Engineering, Architecture, Development, SRE, and Operations teams.
Inclusive workplace culture focused on diversity, equity, and professional growth.
Position specifically designed as an affirmative opportunity for women, supporting greater gender equity in technology leadership.
Access to initiatives and communities focused on women's development, leadership, networking, and career advancement.
Remote work arrangement in Brazil.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
Sourced from lever · view original
Let the agent run this one for you.
Tailored resume, auto-apply, and referral lookup — in under 2 minutes.
Section · 02
Skills
Section · Company
About Jobgether
Jobgether
About