Observability engineer
Location: Stockholm, Sweden
Contract Type: Permanent
Job Reference ID: JR103122
ApplyOur DevOps culture is strong: we believe in “release small, release often.” Thanks to our fully automated CI/CD pipelines, we push over 13,000 releases to production every year, delivering value to customers in real time. We use modern technologies across backend, frontend, apps, data and infrastructure, and we’re continuously evolving. We tackle complex engineering challenges such as real-time scalability, database migrations with no downtime, and peak transaction loads, all while maintaining an exceptional and safe customer experience.
Turning skills into thrills
- Technical Expertise
- Strategic Thinking and Innovative
- Adaptive Problem-Solving & Experimentation
- Cross-Functional Collaboration
- Agility & Adaptability
- Curiosity & Continuous Learning
Turning skills into thrills
Culture points
We take pride in what we do and strive for excellence. We embrace challenges, take initiative, and push boundaries to achieve the best results. With a strong performance-driven mindset, we focus on creating value and continuously aim higher.
We take ownership of our goals and understand our role in delivering results. We seek feedback, take responsibility from start to finish, and hold ourselves to high standards. With a mindset of continuous improvement, we challenge ourselves to grow and do better every day.
We are one team, united by a shared purpose. We work together, support each other, and put the Group’s success first. We embrace diversity, stay open-minded, and face challenges with courage and care. By trusting and lifting one another, we create an environment where we can all thrive.
Observability engineer
At FDJ UNITED, we don't just follow the game, we reinvent it.
FDJ UNITED is one of Europe’s leading betting and gaming operators, with a vast portfolio of iconic brands and a reputation for technological excellence. With more than 5,000 employees and a presence in around fifteen regulated markets, the Group offers a diversified, responsible range of games, both under exclusive rights and open to competition. We set new standards, proving that entertainment and safety can go hand in hand. Here, you’ll work alongside a team of passionate individuals dedicated to delivering the best and safest entertaining experiences for our customers every day.
We’re looking for bold people who are eager to succeed and ready to level-up the game. If you thrive on innovation, embrace challenges, and want to make a real impact at all levels, FDJ UNITED is your playing field.
Join us in shaping the future of gaming. Are you ready to LEVEL-UP THE GAME?
We are looking for a talented and driven Site Reliability Engineer (Observability Engineer) to join our team and play a key role in building the next generation of our high-performance observability framework.In this role, you will be at the heart of designing, developing, and operating the systems that give our engineering organisation deep visibility into the health, performance, and reliability of our products and infrastructure. You will work closely with engineering teams across the business to ensure our observability stack is scalable, reliable, and fit for purpose in a fast-moving, high-stakes environment.
The current focus of the framework is built around VictoriaLogs and VictoriaMetrics, with a strong emphasis on OpenTelemetry for metrics and log ingestion. You will help shape how this framework evolves and be a hands-on contributor to its ongoing development, as well as help maintain this in a high traffic production site.
What You'll Be Doing
Designing and developing a next-generation, high-performance observability platform using modern tooling and best practices.
Driving the adoption and implementation of OpenTelemetry (OTEL/OTLP) for metrics and log ingestion across the organisation.
Building and maintaining Infrastructure as Code (IaC) and configuration management pipelines to automate platform operations.
Contributing to CI/CD pipelines and deployment workflows to support continuous delivery of observability components.
Owning and improving the reliability, scalability, and performance of the observability stack.
Collaborating with engineering teams to define and track Service Level Indicators (SLIs) and establish reliability standards.
Participating in code reviews, architectural discussions, and technical planning sessions.
Supporting capacity planning, release processes, and change management practices.
Championing an automation-first mindset and contributing to a culture of operational excellence.
What We're Looking For
Infrastructure & Configuration Management
Proven experience with Infrastructure as Code, specifically Terraform
Configuration management experience using Ansible and/or Chef
Coding Skills
Strong scripting and development capabilities in Bash, Python, and Go
Experience with Rust and/or Java is a nice to have
CI/CD & Version Control
Solid understanding and hands-on experience with Git and GitLab
Experience with ArgoCD and/or Jenkins for CI/CD pipeline management
Monitoring & Observability
Metrics platforms: Prometheus, Mimir, InfluxDB, VictoriaMetrics
Log aggregation: Splunk, Loki, Elasticsearch, VictoriaLogs
Visualisation: Grafana
Log forwarding: FluentBit, Vector.dev
OpenTelemetry (OTEL/OTLP): Metrics and log pipeline experience
Containers & Orchestration
Strong experience with Docker and/or Podman
Kubernetes (including workload management and cluster operations)
Helm for Kubernetes package management
Cloud
Hands-on experience with AWS
Core Concepts & Principles
You'll need a solid understanding and proven experience of the following:
Systems Architecture: ability to design and reason about complex distributed systems
Networking principles: TCP/IP, DNS, load balancing, proxies, and beyond
TSDB (Time Series Database) principles: understanding of storage, retention, cardinality, and query patterns
Logging principles: structured logging, log pipelines, and log management at scale
Messaging queues: pub/sub patterns, event-driven architectures
Clustering principles: high availability, fault tolerance, distributed coordination
Load balancing: strategies, implementations, and traffic management
API / JSON principles: RESTful APIs, data serialisation, and integration patterns
Caching mechanisms: In-memory caching, external caches, and cache invalidation strategies
Code reviews: Contributing to and leading thoughtful, constructive reviews
Testing and documentation: Writing quality tests and clear technical documentation
Agile mindset: Comfortable working in iterative, collaborative delivery environments
Automation mindset: Defaulting to automation and repeatability in everything you build
Nice to Have
Experience with AI concepts and tooling, including the creation of MCP's and agent Skills
Microservices architecture: Design patterns, service meshes, inter-service communication
Operational Experience
Change Management: Structured approaches to managing infrastructure and service changes
Service Level Indicators (SLIs): Defining, measuring, and acting on reliability metrics
Capacity Planning: Forecasting infrastructure needs based on usage trends and growth
Release Processes: Experience with progressive delivery, feature flags, and rollback strategies
We believe talent knows no boundaries. Our hiring process focuses solely on your skills, experience, and potential to contribute to our team. We welcome applicants from all backgrounds and evaluate each candidate based on merit, regardless of personal characteristics as the age, gender, origin, religion, sexual orientation, neurodiversity or disability.
Testimonials
-
Bertrand Le Piolot
My mission is to position cybersecurity as a business enabler, by finding the right balance between security requirements and business development objectives.
Bertrand Le Piolot
Group Cybersecurity Director -
Lesya Liskevych
Our team turns every customers interaction into mainingful insights, leveraging AI to personalise and enhance the user experience on our gaming platform.
Lesya Liskevych
Head of Product Insights & AI Automation Technology -
From improving product features to enhancing safe gaming practices, data isn't just information, it's a catalyst for innovation and maintaining the Group's integrty.
Nonna Shakhova
Cloud Data EngineerNonna Shakhova
A European gaming champion
FDJ UNITED is a European leader in betting and gaming, trusted for its iconic brands and technological strength across around 15 regulated markets. We’re rapidly digitising our lottery business and expanding our sports-betting footprint, creating exciting opportunities to build the next generation of player experiences. Here you’ll work on high-impact projects: modernising platforms, scaling data-driven personalisation, and developing tools that both delight customers and protect them. Our goal is to strengthen customer relationships through smarter identification and insights. That means meaningful, purpose-driven work, from customer service to marketing, product design, compliance and more. All within an international, innovation-focused environment. We are shaping the future of gaming, join us!
Benefits
Challenge your thoughts
-
-
-
-
Tech is the driving force behind everything we do.
-
We are driving gaming and betting forward with purpose and positive impact.
-
Discover our latest videos
-
FDJ UNITED employees are proud to work with us, thanks to stimulating professional challenges and attention to quality of life at work.
-
Découvrez la déclaration d’accessibilité de FDJ UNITED, notre niveau de conformité et les actions mises en œuvre pour améliorer l’accessibilité de nos services numériques.
-
LET’S STAY IN TOUCH
Don't see what you are looking for? Sign up and we'll notify you when roles become available.