Linux Infrastructure Engineer
Uvation
Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
- Strong understanding of server hardware, including:
- BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
- NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting
- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
- Understanding of NVIDIA GPU technologies including:
- A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization
- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
- Understanding of:
- GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design
- Familiarity with NVIDIA ecosystem technologies such as:
- CUDA
- NCCL
- GPUDirect Storage
- NVIDIA Fabric Manager
- NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
- Advanced Linux storage administration:
- LVM
- XFS, EXT4
- NFS
- iSCSI
- Fibre Channel SAN
- Multipath I/O
- Strong hands-on experience with Ceph , including:
- Cluster architecture
- MON, OSD, MDS
- RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery
- Experience with high-performance AI storage platforms such as:
- WEKA
- VAST Data
- Dell PowerScale
- Pure Storage FlashBlade
- NetApp
- Understanding of:
- NVMe-over-Fabrics (NVMe-oF)
- RDMA
- GPUDirect Storage
- Parallel file systems
- AI data pipelines
Networking & Infrastructure
- Strong networking knowledge:
- Bonding
- VLANs
- Routing
- MTU optimization
- DNS
- DHCP
- Experience with high-performance data center networking:
- 100G/200G/400G Ethernet
- RoCE
- RDMA
- Spine-Leaf architectures
- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
- Experience with high availability, clustering, and disaster recovery
- Strong troubleshooting skills across:
- Linux operating systems
- Hardware platforms
- GPU infrastructure
- Networking
- Enterprise storage
- Experience supporting mission-critical production environments
- Bash and Python scripting for automation and operational efficiency
- Experience creating operational documentation, runbooks, and infrastructure standards
- Understanding of AI infrastructure design and reference architectures
- AI cloud integration for workloads
- SOP and runbook development and maintenance
- Incident, problem, and capacity management
- Business continuity and disaster recovery planning for AI workloads
- Proactive risk identification and mitigation to avoid business impact
Nice to Have
- Kubernetes infrastructure (especially AI/ML and GPU integration)
- KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
- Ansible automation
- NVIDIA Base Command Manager
- Slurm or HPC workload schedulers
- Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
- Data Center Infrastructure Management (DCIM) tools
- IPAM solutions
- AWS, Azure, or hybrid cloud exposure
We Are Not Looking For
- Candidates whose experience is primarily CI/CD pipeline engineering
- Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
- Cloud-only administrators with limited bare metal, storage, or hardware experience
- Professionals whose primary expertise is software development rather than infrastructure engineering
Ideal Candidate
Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.
- ...communication. Preferred: experience optimizing UI rendering and memory management and building applications for macOS, Windows, Linux, or other platforms. Bonus: contributions to open-source projects, especially in P2P or decentralized technology. Benefits ~...
- ...person, we’d love to talk to you. THE ROLE Integrate security into BJAK’s software delivery and infrastructure workflows. Work with Platform, Cloud, and Engineering to automate controls and support secure delivery across AWS and GCP. WHAT YOU WILL BUILD –...
- ...to technical and non-technical stakeholders. Partner with engineering teams, share security knowledge, and improve security... ...applications. Experience in cloud security architecture and infrastructure. Experience coding in Java, Python, or Go, plus at least...
- ...integrations. Improve developer tooling, Supabase services, and infrastructure supporting Next.js and React applications. Troubleshoot... ...At least 3 years of DevOps or infrastructure engineering experience focused on mobile and web platforms. Strong hands...
- ...applications and next steps. Our partner is looking for a Senior DevOps Engineer based in Saudi Arabia. Join an international banking project... ...development and platform teams to provision and manage cloud infrastructure efficiently. The role combines AWS, Terraform, Infrastructure...
$1000 per year
...for the Camunda SaaS offering. Collaborate with product, engineering, and field teams to define, ship, and iterate on features across... ...preferred. Familiarity with multi-region or multi-cloud infrastructure and Crossplane is preferred. Benefits Fully remote and...- ...building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global... ...internally. – Own the governed integration gateway as an engineering surface. You’ll handle connector onboarding, authentication...
- ...manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Saudi Arabia. This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You...
- ...and next steps. Our partner is looking for a Senior Software Engineer, Quality based in Saudi Arabia. This role focuses on building... ...triage. The role combines hands-on software engineering with test infrastructure, distributed systems, CI at scale, and emerging LLM...
- ...About the Role We are looking for a highly skilled Frontend Engineer (React.js) with 4–7years of professional experience, ideally gained in fast-paced startup environments. This role requires strong expertise in building modern, scalable, and high-performance user...Remote job
- ...partner is looking for a Enterprise AI Governance & Trust Layer Engineer based in Saudi Arabia. This is a remote engineering role... ...experienced engineer who wants to shape secure and responsible AI infrastructure. Accountabilities - Design and deploy enterprise AI trust...
- ...time user experiences to millions of active users by delivering high-throughput, low-latency Flutter applications while elevating engineering standards across our development teams. What you’ll drive: Strategy & outcomes Own high-level mobile system architecture...
- Responsibilities Deliver new features and bug fixes with an emphasis on reliability, performance, and user satisfaction. Design and develop responsive, high-performance user interfaces and application state across platforms. Maintain code quality through coding...
- ...Conduct threat modeling and design reviews for features, APIs, third-party integrations, and major changes. – Coordinate with mobile engineers on cross-platform findings and backend controls protecting native iOS and Android clients. – Own application security...
- ...applications and next steps. Our partner is looking for a Senior Data Engineer (Python / AWS / ML Pipelines) based in Saudi Arabia. As a... ...environment. You will work across data engineering, cloud infrastructure, and machine learning operations to turn models into reliable...
- ...We are seeking a Senior iOS Engineer to own the client side implementation of how members join, pay, and stay with Raya. Member Experience owns the full member lifecycle – applying, onboarding, payments, and lifecycle management – and this role owns the surfaces where...
- ...other languages as needed. Debug and resolve performance issues using experiments, measurements, and profiling tools. Educate engineers through code reviews, talks, tooling, and documentation. Define and drive actionable next steps from high-level objectives....
- ...is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Saudi Arabia. You will work at the intersection of AI research, systems engineering, and...
- ...data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's...
- ...countries around the world. GE Vernova’s Gas Power business engineers advanced, efficient natural gas-powered technologies and... ...generation equipment, enabling operators of the world’s energy infrastructure to provide more reliable and affordable energy. Job...Remote job
- ...countries around the world. GE Vernova’s Gas Power business engineers advanced, efficient natural gas-powered technologies and... ...generation equipment, enabling operators of the world’s energy infrastructure to provide more reliable and affordable energy. Job...Remote job
- ...management ecosystem. Collaborate with DataHub's product and engineering teams to understand technical concepts deeply and translate... ...Tackle high-impact challenges at the heart of enterprise AI infrastructure Ship production systems that power real-world use cases...
- ...multi-workstream engagement across the region's sovereign AI infrastructure landscape. This is a high-visibility, single-account-depth... ...Bring partner feedback and requirements back to product and engineering, and advocate for the resourcing needed to keep commitments...
- ...you will shape enterprise security strategies across cloud, infrastructure, identity, and data environments. You will define current-state... ...architecture role. - Bachelor’s degree in Computer Science, Engineering, Cybersecurity, or a related discipline . - Strong expertise...
- ...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes... ...to design, deploy, and optimize production-grade AI infrastructure and agent systems. You’ll be responsible for architecting scalable...
- ...countries around the world. GE Vernova’s Gas Power business engineers advanced, efficient natural gas-powered technologies and... ...generation equipment, enabling operators of the world’s energy infrastructure to provide more reliable and affordable energy. Job Summary...Remote job
- ...countries around the world. GE Vernova’s Gas Power business engineers advanced, efficient natural gas-powered technologies and... ...generation equipment, enabling operators of the world’s energy infrastructure to provide more reliable and affordable energy. Job...Remote job
- ...Coding Specialist and AI Trainer, you’ll use your software engineering expertise and Dutch fluency to help improve the capabilities... ...programming challenges across algorithms, architecture, development, infrastructure, and systems programming. The role combines hands-on...
- ...About the Company Armada is the hyperscaler for the edge, delivering modular AI infrastructure from first deployment to AI factory with speed, scale and sovereignty. Named one of Fast Company's Most Innovative Companies and to the CNBC Disruptor 50, Armada’s solutions...
- ...-visibility role requires deep expertise in BambooHR and modern HR technology suites. You will take full ownership of our HR infrastructure: managing candidate pipelines, screening resumes with sharp discernment, drafting offer letters, maintaining compliance, and providing...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Linux Infrastructure Engineer. Be the first to apply!
