Dheeraj Ramasahayam

I build datacenter networks that fix themselves.

Network engineer with 3+ years in hyperscale data centers, and researcher in AI-driven failure prediction, root cause analysis, and autonomous recovery.

Live above: a Clos fabric losing a link, rerouting, and healing — my research automates exactly this: predict, detect, localize, recover.

The Problem

Every AI model, hospital record, financial transaction, and video call crosses a datacenter network. When those networks fail, recovery is still largely manual: engineers paging through dashboards while services degrade. Hyperscale operators measure these outages in millions of dollars per hour — and the AI buildout is multiplying network scale faster than operations teams can grow. My work attacks this from both ends: hands-on infrastructure delivery for the largest cloud operators, and published research toward networks that predict, diagnose, and repair their own failures.

Research

Seven papers forming one program: predict failures → detect intrusions → benchmark root cause analysis → unify into a self-healing system → make it explainable → federate it across sites → harden the wireless edge.

  1. Published · Authorea · Mar 2026

    Temporal Attention-Guided Sequence Learning for Zero-Shot Generalization in Datacenter Network Failures

    An early-warning system that predicts datacenter network failures up to 60 seconds before they happen. Combines an LSTM sequence model with a custom self-attention mechanism, and generalizes zero-shot to failure types it was never trained on.

  2. Published · Authorea · Mar 2026

    Drift-Adaptive Intrusion Detection for Enterprise Networks

    Enterprise intrusion detection that keeps working as attack traffic evolves. Aligns four public benchmarks (UNSW-NB15, NSL-KDD, CICIDS2017, CSE-CIC-IDS2018) into one 41-feature representation and adds an online drift-adaptive controller on top of a hybrid ensemble.

  3. Published · Authorea · Mar 2026

    ClosRCA-Bench: An Open Topology-Grounded Benchmark and Counterfactual Recovery Framework for Self-Healing Datacenter Networks

    An open benchmark for root cause analysis in Clos datacenter fabrics, with fixed splits, public scripts, and a counterfactual framework that validates whether a proposed fix would actually restore the network. Dataset and code archived on Zenodo (DOI 10.5281/zenodo.19059194).

  4. Published · Authorea · Mar 2026

    Autonomous Self-Healing Datacenter Networks: A Unified AI System for Prediction, Detection, Root Cause Analysis, and Recovery

    The flagship systems paper: a unified AI pipeline that predicts failures from telemetry, detects intrusions under drift, localizes root causes on the topology, and executes safety-gated recovery validated in a digital twin before touching production.

  5. Published · Authorea · Mar 2026

    Explainable AI for Root Cause Analysis in Large-Scale Datacenter Networks

    Makes AI-driven root cause analysis trustworthy for operators: explainability methods that show why the model blames a device or link, evaluated for scalability on large datacenter topologies.

  6. Preprint · Under review

    Federated and Continual Learning for Autonomous Self-Healing Datacenter Networks

    Extends self-healing networks across organizations: federated and continual learning so multiple datacenters improve a shared model without sharing raw telemetry, and without forgetting older failure modes.

  7. In preparation

    Self-Evolving AI Framework for Evaluating Wireless Authentication Robustness in WPA2/WPA3 Networks

    A self-evolving AI framework for stress-testing WPA2/WPA3 wireless authentication, probing robustness of enterprise Wi-Fi security. Manuscript in preparation.

Patent — Self-Evolving Multi-Reality Network Infrastructure

U.S. Provisional Application No. 64/020,128 · Filed March 28, 2026 · Patent Pending

An infrastructure control system that maintains multiple candidate 'realities' of a network — simulated execution branches with temporal state rewriting — and recursively compares their predicted outcomes against live observations. The system evolves its own control policies and commits only changes proven safe in simulation, extending self-healing from reaction to anticipation.

Field Work

Deployment log — hyperscale and enterprise infrastructure delivered.

Projects & Code

clos-rca-bench

Open RCA benchmark with Zenodo-archived dataset, fixed splits, CI, and reproduction scripts.

xai-network-rca

Explainable RCA toolkit; IEEE TNSM draft, tested Python package with CI.

ai-network-threat-detection

Drift-adaptive IDS benchmark across four public datasets; Dockerized, reproducible.

datacenter-network-failure-predictor

PyTorch LSTM+attention failure predictor with notebooks, configs, and results.

Enterprise Monitoring & SIEM Lab

Proxmox-based enterprise lab with SIEM alerting, firewall log integration, and Azure monitoring deployment (40% faster incident response).

Virtualized Home Infrastructure

Multi-VM Proxmox environment for VPN, media, and Linux services; 30% hosting cost reduction.

Skills & Certifications

Networking & Protocols

Hardware & Optical

Monitoring & Analysis

Cloud & Virtualization

Scripting & ML

CCNA · CEH v11 (EC-Council) · IBM Cybersecurity Analyst · CompTIA A+ · CompTIA ITF+ · Certified Scrum Master · OSHA 10

M.S. Computer Information Systems, New England College (2024) · B.E. Electrical & Electronics Engineering, NIT Andhra Pradesh (2022)

Contact

dheerajramasahayam@gmail.com · GitHub · LinkedIn · Resume (PDF)