Rojan Adhikari

Rojan Adhikari

Data Engineer

Cloud Data Platforms · Streaming Systems · Infrastructure Automation

I build reliable, scalable data platforms using AWS, Apache Flink, Kafka, Terraform, and Kubernetes.

I work at the intersection of data engineering, cloud infrastructure, and platform reliability.

Open to new opportunities

Last updated: July 2026

About

I'm a data engineer specializing in cloud data platforms, infrastructure automation, and real-time data processing. My work involves designing and supporting systems built with AWS, Apache Flink, Kafka, Kubernetes, Terraform, and Python.

I currently contribute to enterprise data platforms where I help build deployment automation, streaming pipelines, infrastructure-as-code, monitoring, access controls, and operational tooling. I also troubleshoot distributed-system issues involving container deployments, IAM permissions, stream processing, data consistency, and pipeline reliability.

I enjoy working on problems where data engineering, cloud infrastructure, and platform reliability come together. My goal is to create systems that are scalable, observable, secure, and easier for engineering teams to operate.

Currently

Location
Based in Dallas, Texas
Current focus
Working on cloud data-platform engineering
Interested in
Interested in Data Engineering, Cloud Data Engineering, Data Platform Engineering, Infrastructure Engineering opportunities

Experience

  1. November 2025 – Present

    Senior Data Engineer

    Duke Energy (Energy and Utilities) · Contract

    Build and operate AWS cloud data infrastructure for an energy and utilities client, provisioning environments with Terraform and Kubernetes while supporting real-time streaming pipelines built on Apache Flink, Kafka, and Debezium.

    • Build and maintain cloud data infrastructure using AWS, Terraform, Kubernetes, Helm, and GitHub Actions.
    • Support streaming-data workloads using Apache Flink, Apache Kafka, Debezium, and Apache Hudi.
    • Develop and improve CI/CD workflows for building, publishing, and deploying containerized applications.
    • Implement monitoring and alerting using Prometheus, Grafana, Grafana Alloy, and Mimir.
    • Investigate incidents involving container deployments, image availability, checkpoint failures, data duplication, permissions, and pipeline performance.
    • Configure AWS services including IAM, Lake Formation, Glue, S3, KMS, SNS, Lambda, EventBridge, Athena, EKS, and Secrets Manager.
    • Automate infrastructure provisioning and environment configuration through Terraform and deployment workflows.
    • Collaborate with application teams, infrastructure engineers, and product teams to troubleshoot platform issues.
    • Participate in production deployments, on-call rotations, root-cause investigations, and platform support.
    • Create documentation and operational runbooks that help teams safely manage cloud resources and data-platform operations.
    • AWS
    • Python
    • Terraform
    • Apache Flink
    • Apache Kafka
    • Debezium
    • Apache Hudi
    • Kubernetes
    • Helm
    • Docker
    • GitHub Actions
    • Grafana
    • Prometheus
  2. Oct 2024 – Oct 2025

    Senior Data Engineer

    Mid First Bank (Financial) · Contract

    Led modernization of a banking data platform, migrating legacy SQL Server and Oracle pipelines to an Azure-based lakehouse on Databricks and Delta Lake, with Airflow-orchestrated ingestion and Snowflake dimensional models for risk, lending, and regulatory reporting.

    • Contributed to the modernization of the bank’s data platform by migrating legacy SQL Server and Oracle data pipelines to an Azure-based lakehouse architecture using Azure Data Lake Storage, Databricks, Delta Lake, and Snowflake.
    • Designed batch and incremental ingestion pipelines using Azure Data Factory, PySpark, and SQL to process customers, account, transaction, loan, and operational data from multiple banking systems.
    • Developed bronze, silver, and gold data layers in Delta Lake, applying schema validation, deduplication, data standardization, and business transformations before publishing curated datasets for analytics.
    • Implemented change data capture and incremental-loading patterns using source timestamps, merge operations, and Delta Lake upserts, reducing unnecessary full-table processing and improving data freshness.
    • Built PySpark and SQL transformation jobs in Databricks to cleanse, join, and aggregate high-volume financial datasets used by risk, lending, finance, and regulatory-reporting teams.
    • Developed dimensional data models in Snowflake, including fact and dimension tables with Slowly Changing Dimension Type 2 logic, to support historical reporting and customer and account analytics.
    • Created and maintained Apache Airflow workflows to orchestrate ingestion, transformation, data-quality validation, and downstream publishing processes with retries, dependencies, and failure notifications.
    • Implemented automated data-quality checks for schema conformity, null values, duplicate records, referential integrity, and source-to-target reconciliation, helping identify data issues before they reached reporting layers.
    • Applied banking-data security controls using Azure RBAC, managed identities, Key Vault, Snowflake role-based access, encryption, and least-privilege permissions to protect sensitive customer and financial information.
    • Developed reusable Terraform modules and GitHub Actions workflows to provision cloud resources, validate infrastructure changes, and promote data-pipeline configurations across development, testing, and production environments.
    • Monitored pipeline execution, Spark workloads, data freshness, and job failures using Azure Monitor, Databricks logs, and operational dashboards; investigated production incidents and documented root causes and corrective actions.
    • Partnered with business analysts, data architects, application teams, and governance stakeholders to translate reporting and data requirements into maintainable data models, pipelines, and technical documentation.
    • Participated in Agile ceremonies, design discussions, pull-request reviews, release planning, and production deployments while supporting and mentoring engineers on PySpark, SQL, and pipeline-development practices.
    • Azure
    • Azure Data Factory
    • Azure Data Lake Storage
    • Databricks
    • Delta Lake
    • Snowflake
    • SQL Server
    • PySpark
    • SQL
    • Apache Airflow
    • Terraform
    • GitHub Actions
    • Azure Monitor
  3. Sept 2023 – Oct 2024

    Data Engineer

    Elara Caring (Healthcare) · Contract

    Built a HIPAA-compliant healthcare data platform on Azure, developing Databricks and PySpark pipelines that ingest and transform patient, claims, and HL7/FHIR clinical data into governed Delta Lake and Synapse models with Purview-tracked lineage.

    • Contributed to the development of a HIPAA - compliant healthcare data platform using Azure Data Factory, Azure Data Lake Storage, Databricks, Delta Lake, and Azure Synapse Analytics.
    • Built batch and incremental ingestion pipelines to integrate patient, clinical, claims, provider, and operational data from EHR systems, relational databases, APIs, and secure file transfers.
    • Developed PySpark transformation jobs in Azure Databricks to cleanse, standardize, deduplicate, and enrich healthcare data before publishing curated datasets for analytics and reporting.
    • Designed Bronze, Silver, and Gold data layers in Delta Lake, separating raw clinical data from validated and business - ready datasets while supporting schema evolution and historical processing.
    • Parsed and transformed HL7 and FHIR healthcare messages using Python, PySpark, and JSON processing, converting nested clinical records into standardized structures for downstream consumption.
    • Implemented incremental data - processing patterns using watermark columns, change timestamps, and Delta Lake merge operations to improve data freshness and avoid unnecessary full-table reloads.
    • Created dimensional data models in Azure Synapse for patient outcomes, episodes of care, provider performance, and operational reporting, including historical tracking for changing patient and provider attributes.
    • Developed and maintained Azure Data Factory workflows to orchestrate ingestion, Databricks transformations, data - quality validation, and downstream publishing with dependency management, retries, and failure notifications.
    • Implemented data - quality checks for required fields, duplicate patient records, invalid clinical codes, schema changes, referential integrity, and source - to - target reconciliation before data reached reporting systems.
    • Protected sensitive patient information through Azure RBAC, managed identities, Key Vault, encryption, private endpoints, and role - based access controls aligned with HIPAA and least - privilege requirements.
    • Implemented metadata management and data - lineage capabilities using Microsoft Purview, helping teams understand the movement and transformation of sensitive healthcare data across the platform.
    • Monitored pipeline execution, data freshness, Databricks workloads, and processing failures using Azure Monitor, Log Analytics, and operational dashboards; investigated incidents and documented root causes.
    • Used Terraform and CI / CD workflows to manage cloud resources and deploy notebooks, pipeline configurations, and infrastructure changes consistently across development, testing, and production environments.
    • Collaborated with clinical analysts, application teams, data architects, security teams, and business stakeholders to translate healthcare reporting requirements into scalable pipelines and trusted datasets.
    • Azure
    • Azure Data Factory
    • Azure Data Lake Storage
    • Databricks
    • Delta Lake
    • Azure Synapse Analytics
    • PySpark
    • Python
    • Microsoft Purview
    • Azure Monitor
    • Terraform
  4. Feb 2021 – July 2023

    Data Engineer

    Humana (Healthcare Insurance) · Contract

    Built an AWS-based healthcare data platform for claims, enrollment, and provider data, using Glue and EMR for PySpark transformations, Airflow-orchestrated ingestion, and Redshift dimensional models with SCD Type 2, while enforcing HIPAA-aligned access controls.

    • Contributed to the development of an AWS - based healthcare data platform supporting claims, member enrollment, provider, authorization, and clinical reporting workloads.
    • Built batch ingestion pipelines using AWS Glue, Python, and SQL to process data from relational databases, partner file feeds, APIs, and application-generated files into Amazon S3.
    • Developed PySpark transformation jobs using AWS Glue and EMR to cleanse, standardize, deduplicate, and enrich healthcare data before publishing curated Parquet datasets for downstream analytics.
    • Organized datasets into raw, validated, and curated S3 layers with consistent partitioning, naming conventions, and retention policies, improving maintainability and query efficiency.
    • Implemented incremental data - loading patterns using processing timestamps, source-system keys, and change indicators to reduce repeated processing of historical claims and enrollment data.
    • Designed fact and dimension tables in Amazon Redshift for claims, member enrollment, providers, and authorizations, including Slowly Changing Dimension Type 2 logic for historical attribute tracking.
    • Developed and maintained Apache Airflow workflows to orchestrate ingestion, transformation, data - quality validation, and Redshift loading processes with dependencies, retries, and operational notifications.
    • Implemented data - quality checks for duplicate claims, missing member identifiers, invalid dates, code - format violations, referential integrity, and source - to - target record reconciliation.
    • Optimized S3 and Redshift workloads through partition pruning, appropriate file sizing, compression, distribution styles, and sort keys to improve recurring analytics and reporting performance.
    • Applied security controls using AWS IAM roles, KMS encryption, S3 bucket policies, Secrets Manager, and role-based Redshift access to protect member and claims information in accordance with HIPAA requirements.
    • Created CloudWatch logs, metrics, and alerts for failed jobs, delayed data, and processing errors; investigated production issues and documented remediation steps for recurring failures.
    • Used Terraform and CI / CD pipelines to provision AWS resources and deploy Glue jobs, Airflow configurations, and supporting infrastructure consistently across environments.
    • Developed curated datasets and SQL views used by Tableau and Power BI dashboards for claims utilization, member enrollment, provider performance, and operational reporting.
    • Collaborated with analysts, application teams, healthcare - domain specialists, and senior engineers to clarify source - system behavior, validate business rules, and deliver reliable datasets for reporting.
    • AWS
    • AWS Glue
    • EMR
    • S3
    • Redshift
    • PySpark
    • Python
    • SQL
    • Apache Airflow
    • Terraform
    • CloudWatch
    • Tableau
    • Power BI

Featured Projects

Real-Time Streaming Data Platform

A cloud-native streaming pipeline that ingests events from Kafka, processes them with Apache Flink, and writes curated data to a cloud data lake.

  • Apache Kafka
  • Apache Flink
  • Python
  • AWS S3
  • Apache Hudi
  • Terraform
  • Grafana
  • Docker

Data Platform Observability

A monitoring and alerting solution for distributed data-processing workloads using Prometheus-compatible metrics and Grafana.

  • Grafana
  • Prometheus
  • Grafana Alloy
  • Mimir
  • Kubernetes
  • Apache Flink
  • Apache Kafka

Data Engineering Operations Toolkit

A Python command-line toolkit for validating cloud resources, checking deployment configuration, and simplifying common data-platform operational tasks.

  • Python
  • AWS SDK for Python
  • Pytest
  • Docker
  • GitHub Actions

Skills

Cloud Platforms

  • AWS
  • Azure
  • GCP

AWS Services

  • S3
  • Lake Formation
  • Glue
  • Glue Data Catalog
  • Athena
  • API Gateway
  • Lambda
  • Step Functions
  • EventBridge
  • IAM
  • KMS
  • Secrets Manager
  • EMR
  • EKS
  • Redshift
  • RDS
  • CloudWatch
  • CloudTrail
  • SNS
  • SQS
  • Kinesis Data Streams

Programming Languages and Libraries

  • Python
  • SQL
  • T-SQL
  • PL/SQL
  • PySpark
  • Bash
  • Pandas
  • Java
  • YAML
  • Markdown

Big Data and Streaming

  • Apache Flink
  • Apache Kafka
  • Apache Spark
  • Spark Structured Streaming
  • Apache Hudi
  • Apache Iceberg
  • Debezium
  • Delta Lake
  • Delta Sharing
  • Change Data Capture
  • Batch Processing
  • Stream Processing
  • Kinesis Data Streams
  • ETL and ELT
  • Data Lakes

Data Orchestration and Workflow Management

  • Apache Airflow
  • dbt
  • AWS Step Functions

Data Architecture and Modeling

  • Medallion Architecture
  • Star Schema
  • Snowflake Schema
  • Dimensional Modeling
  • Slowly Changing Dimensions (SCD Type 2)
  • Parquet
  • JSON

Databases and Data Warehouses

  • Snowflake
  • Amazon Redshift
  • PostgreSQL
  • SQL Server
  • MongoDB

Data Quality and Observability

  • Grafana
  • Prometheus
  • Grafana Alloy
  • Mimir
  • CloudWatch
  • Azure Monitor
  • Datadog
  • Great Expectations
  • PyTest
  • Data Profiling
  • Root-Cause Analysis
  • Alerting

Security and Compliance

  • IAM
  • AWS Lake Formation
  • KMS
  • Vault
  • AWS Secrets Manager
  • OAuth2
  • RBAC
  • TLS/SSL
  • HIPAA
  • PCI-DSS
  • GDPR

Infrastructure as Code and DevOps

  • Terraform Enterprise
  • Docker
  • Kubernetes
  • AKS
  • EKS
  • Helm Charts
  • GitHub Actions
  • Jenkins
  • CI/CD Pipelines

Data Cataloging and Governance

  • Azure Purview
  • AWS Glue Data Catalog
  • Data Lineage Tracking
  • Metadata Management
  • Data Governance

MLOps and AI

  • MLflow
  • SageMaker Feature Store
  • Vertex AI
  • EvidentlyAI
  • Data Contracts
  • Feature Pipelines
  • DataOps
  • FastAPI
  • Model Monitoring
  • Model Drift Detection

Business Intelligence and Visualization

  • Tableau
  • Power BI
  • Looker
  • Looker Studio

Collaboration and Project Management

  • Git
  • GitHub
  • Jira
  • Confluence
  • Agile/Scrum
  • Distributed Systems Troubleshooting
  • Production Support
  • Technical Documentation

Education

May 2025

Master of Science in Computational Science – Computer Science

University of Central Oklahoma - Edmond, Oklahoma

Graduate study focused on computer science, computational methods, software development, data processing, and applied technical problem-solving.

Relevant Coursework

  • Distributed Systems
  • Cloud Computing
  • Database Systems
  • Data Structures and Algorithms
  • Software Engineering
  • Machine Learning
  • High-Performance Computing

Contact

I’m interested in opportunities involving data engineering, cloud infrastructure, streaming systems, and platform engineering. Feel free to reach out if you would like to discuss a role, project, or technical collaboration.

rojan.adhikari23@gmail.com
Location
Dallas, Texas