Technical Monitoring & Integration Engineer

The Technical Monitoring & Integration Engineer is responsible for designing, configuring, and maintaining monitoring solutions that support a 24/7 Network Operations Center (NOC). This role focuses on Grafana, Prometheus, Python automation, distributed network monitoring, and integrations with external systems such as Zendesk. The engineer manages multiple Grafana instances in production and ensures reliable observability across large, geographically dispersed environments including retail locations, hotels, and multi site enterprise networks.

What will be your key responsibilities:

Responsibilities 

  • Grafana Architecture Design, configure, and maintain multi tenant Grafana environments, including dashboards, alerting rules, folder structures, and RBAC.
  • Prometheus Configuration Deploy, tune, and maintain Prometheus servers, exporters, and scrape configurations for diverse network and system metrics.
  • Distributed Network Monitoring Implement monitoring for distributed retail and hotel networks, including WAN health, site availability, and device reachability.
  • Fortinet MIB Polling Configure Prometheus SNMP exporters to poll Fortinet firewalls using Fortinet MIBs, collecting metrics such as CPU, sessions, bandwidth, and interface status.
  • Network Equipment Reachability Build monitoring workflows that perform ICMP ping checks, latency measurements, and packet loss detection across customer sites.
  • Python Automation Develop Python scripts to automate monitoring tasks, API integrations, SNMP data processing, and operational workflows supporting NOC functions.
  • Zendesk Integration Build and maintain integrations between Grafana and Zendesk using webhook automation to ensure alerts generate actionable tickets.
  • Multi Instance Management Manage multiple Grafana instances (two per customer), ensuring consistency in configuration, security, and dashboard standards.
  • Monitoring Operations Support NOC teams by ensuring monitoring systems are accurate, stable, and aligned with customer SLAs and operational requirements.
  • Documentation & Standards Produce clear documentation for dashboards, integrations, scripts, SNMP configurations, and operational procedures.

What experience should you have:

Technical expertise:

  • Hands on experience with Grafana in multi instance, multi tenant environments.
  • Strong proficiency with Prometheus, including SNMP exporters, Alertmanager, and performance tuning.
  • Experience monitoring distributed networks such as retail stores, hotels, or multi site enterprise networks.
  • Ability to configure SNMP polling using Fortinet MIBs and interpret firewall metrics.
  • Experience implementing ICMP ping checks, latency monitoring, and network device availability workflows.
  • Solid Python scripting skills for automation, API integrations, and data processing.
  • Experience integrating monitoring systems with Zendesk using webhook automation.
  • Understanding of NOC workflows, incident management, and operational SLAs.
  • Familiarity with Linux environments, containers, and basic networking concepts.

Preferred Qualifications

  • Experience with Grafana Loki, Alertmanager, or other observability tools.
  • Knowledge of SNMPv3 security, OIDs, and multi vendor MIBs.
  • Exposure to MSP or multi customer NOC environments.
  • Experience with CI/CD pipelines for deploying monitoring configurations.

Success Criteria 

  • Reliable monitoring across distributed customer networks including retail and hotel locations.
  • Accurate SNMP polling of Fortinet devices and actionable dashboards for NOC agents.
  • High quality alerting pipelines that reduce noise and improve incident response.
  • Stable and well documented integrations between Grafana, Prometheus, and Zendesk.
  • Automation that reduces manual effort and improves operational efficiency.

What do you get in return:

 Our team is composed of experts in their fields who are passionate about delivering high-quality work and maintaining a positive work culture. We value innovation, teamwork, and personal growth. As an experienced Integration Engineer, you will have the opportunity to make a significant impact on our projects and contribute to the success of our organization. If you are ready to embrace exciting challenges and foster a culture of excellence, we encourage you to apply.

What do we offer:

  • Work remotely from anywhere in the world, with a fully remote team, and enjoy a mutually agreed schedule that fits your needs. (Core US working hours) 
  • Work primarily with US-based colleagues, providing you with the opportunity to collaborate with people from diverse backgrounds and skill sets.
  • Use your skills and expertise to make a significant impact on the delivery of projects in our company
  • Work in a supportive environment that values your contribution and provides you with the resources and training you need to grow in your career.
  • Enjoy a 40-hour workweek that provides you with a healthy work-life balance, and the time to pursue your personal and professional goals outside of work.

Mám zájem o tuto pozici

Poslat nabídku na e-mail

Další pozice v oboru Informační technologie, region remote

Linux Administrator

  • Goodcall
  • Pardubický kraj
  • Dohodou

Pro světového výrobce elektroniky hledáme Linux administratora. Znáte skriptování a SQL databáze? Zajímáte se o IT technologie a jste týmový hráč? Pak čtěte dál!

Linux Administrator

Software Engineer

  • Košík
  • Praha hl.m.
  • Dohodou

SW Engineer je full-stack vývojář, který navrhuje, dodává a provozuje kritický podnikový software od návrhu až po produkční provoz s přístupem AI-first napříč celým životním cyklem vývoje. Tato role…

Software Engineer

Senior Product Owner

  • Aures
  • Praha hl.m.
  • Dohodou

Jsme technologická divize nadnárodní skupiny AURES Holdings. Vyvíjíme e-commerce a retail softwarovou platformu, navazující integrace a pomáháme firmám ve skupině napříč Evropou s digitální a…

Senior Product Owner