Creator Marketing moves fast, but the data behind it has always been fragmented, inconsistent, and built for the moment, not for scale.
Primetag was built to change that. We are the Creator Intelligence Data infrastructure platform: structuring and processing the world's largest Creator & Social Content dataset — 10+ billion pieces of content, +15M added daily — and turning fragmented creator activity into standardized, indexed, enterprise-grade data. Today, over 1000 global brands and 260 agencies trust Primetag as the reliable foundation for their Creator Marketing operations.
And now, that foundation powers something bigger. Universe by Primetag is the AI Operating System for Creator Marketing — an agentic AI layer trained on the world's deepest creator dataset. Less guesswork. More impact. At scale.
This is a full-time, fully remote position with a key role in ensuring the reliability, scalability, and efficiency of Primetag’s production infrastructure.
You will join a team of Backend and Systems Engineers responsible for infrastructure, core services, data pipelines, and our observability platform. Within the team, you will help deliver production-grade services with demanding requirements in terms of performance and cost-effectiveness, while ensuring key aspects such as documentation, testing, and Application Performance Monitoring (APM) metrics.
You will contribute at both the strategic and execution levels, bringing a strong understanding of complex systems and applying it to demanding production environments. You will be encouraged to challenge existing approaches, raise your hand when something doesn’t seem right, and identify better and more effective ways of working.
Our systems process billions of social media data points across 50+ markets, handle millions of API requests per day, and support a platform that clients depend on around the clock.
We are looking for an experienced Site Reliability Engineer with 5+ years of experience who thrives in high-data, high-availability environments and is comfortable working with complex production systems.
You combine strong technical expertise with a sharp eye for design and enjoy working across infrastructure engineering, automation, observability, and reliability. You take ownership, care about how systems perform in production, and are willing to challenge the status quo when you see an opportunity to make things more reliable, scalable, or efficient.
If you've operated production infrastructure at this scale, you'll feel right at home.
Own and evolve our observability stack (LGTM), from alert tuning and dashboards to SLO definition;
Maintain and improve our CI/CD pipelines (GitLab CI) and deployment processes across our Kubernetes clusters;
Ensure platform scalability and cost-efficiency as data volumes and client load continue to grow;
Foster SRE best practices across the Engineering team, including SLOs, error budgets, and runbooks.
More than 5 years of professional experience in SRE, DevOps, or infrastructure engineering;
Strong hands-on experience with Kubernetes (AKS preferred), Helm, Flux, Docker, and Ansible in production environments;
Experience with Microsoft Azure, including managed services such as AKS, Azure Storage, Service Bus, MySQL, PostgreSQL, and CosmosDB;
Proficiency with GitLab CI/CD and GitOps principles;
Hands-on experience with Terraform;
Experience with observability stacks, specifically Grafana, Loki, Tempo, Mimir, and Prometheus, or equivalent technologies to the LGTM stack;
Comfortable working with SRE principles, including SLOs, error budgets, incident management, and on-call processes;
Must be a fiscal resident of Portugal, regardless of nationality.
We can provide more information about our technology stack once you initiate your application.
Knowledge of configuring and managing SQL and NoSQL databases, including MongoDB for large data volumes, as well as search engines such as OpenSearch or Elasticsearch;
Familiarity with event streaming and message queues, including Azure Service Bus, Kafka, or RabbitMQ;
Experience with cloud cost optimization, including identifying waste, rightsizing resources, and reporting on infrastructure spend;
Familiarity with DevSecOps, including vulnerability scanning in CI/CD pipelines, secret management, or security hardening;
Experience with Python, particularly reading and understanding service code in FastAPI or similar frameworks. You won't be expected to write features, but you'll need to instrument and debug services.
Help lead a paradigm shift in how B2B influencer marketing is executed and scaled.
Join a mission-driven, innovation-focused company in a fast-growing sector.
Work with a talented, ambitious, and collaborative international team.
Benefit from flexibility, autonomy, and the opportunity to shape the future of B2B marketing in our industry.
The timing is perfect — join us during an exciting expansion phase, with new products and markets.
Work from anywhere within ±2 hours of the CET time zone and take part in our annual “Nomad Offices” in amazing locations (the last one was in Sardinia, Italy).
Be part of a collaborative and multicultural culture, driven by product excellence and real impact.
Receive a competitive salary package and join a Martech company with the potential to become a future “unicorn”.
Made a quick presentation of yourself;
Met every team to get to know them and understand what they do, with a particular focus on the Engineering team to understand the services architecture and their main reliability challenges;
Shadowed your SRE mentor to understand our current processes for deployments, incident handling, and monitoring;
Made your first Merge Request, improving onboarding documentation and general guidelines.
After your first week, you should feel comfortable navigating the infrastructure and know who to ask for help.
Have a general understanding of the company’s mission, values, and culture, as well as the services we provide;
Take increasing ownership of day-to-day operational tasks, including monitoring, alert handling, and supporting Engineering teams;
Take responsibility for smaller tickets related to infrastructure, CI/CD, or monitoring, under the guidance of your mentor.
After 3 months, you should be able to independently handle routine operational tasks, contribute to infrastructure automation, and support incident response with minimal guidance.
Be comfortable proposing and introducing new techniques and tools for our daily work;
Have contributed to at least one of our OKRs and major goals;
Own specific areas of our reliability stack, such as monitoring, alerting, or CI/CD;
Proactively identify bottlenecks and propose improvements in scalability, security, and resilience;
Operate with confidence and independence in maintaining and improving our production infrastructure.
After 6 months, you should be trusted to fully own critical parts of our infrastructure, operate with autonomy, and actively contribute to setting our long-term reliability strategy.
Submit your application using this form.