NSPNikhil Sai Srinivas Pokuri Start a conversation ↗

Available for the next hard problem

Data platforms
that hold up.

I’m Nikhil — a Senior Data Engineer and Retail Data Platform Lead simplifying complex production systems, owning the messy issues, and leaving teams with reusable ways to move faster.

Ownership

Large, complex Retail segment

Reuse

20+ source pipeline framework

Recovery

About a week of batch runs

Leverage

Two-line prompt workflow

01 / why I stand out

The portfolio is about the operating judgment around the pipeline: how complexity gets reduced, how evidence gets built, and how engineering work becomes easier to repeat.

Production ownership

Not just development — production issues, CRs, recovery, and LOB support.

Platform thinking

Reusable pipelines and simplification of complex Retail systems.

Business + technical

Controls, business requirements, lineage, and stakeholder interaction.

Engineering innovation

AI-powered automation and continuous improvement.

01

Big Data

02

AWS Data Engineering

03

Data Platform Engineering

04

Databricks / Lakehouse

02 / selected work

The useful
mess.

Six pieces of work that show how I approach platform complexity: establish evidence, take the ownership, build the repeatable shape, and keep the business in the picture.

01Production systems · ownershipRetail Platform OwnershipA large and complex Retail segment simplified through technical and operational ownership.

Problem: A complex Retail area required broader ecosystem understanding, technical ownership, and less onshore dependency.

What I did: Selected as Retail Lead; simplified workflows, handled business and LOB issues, supported production, shared knowledge, and helped the team operate more independently.

Technology: AWS · PySpark · Spark · SQL · Data platforms

Outcome: Reduced onshore dependency through clearer ownership and simpler ways of working.

Retail LeadProduction ownershipPlatform simplification
02Production intelligence · controlsProduction Intelligence & Data Control FrameworkAudit data turned into technical KPIs, business controls, and anomaly visibility.

Problem: Pipeline health needed to be visible beyond job completion.

What I did: Implemented audit tables with start and end times, source, target, and reject counts; built runtime and completion KPIs and source-based controls to surface sudden drops and anomalies.

Technology: SQL · Data quality controls · Production metrics

Outcome: Better evidence for operational decisions and business issue investigation.

Data controlsRuntime trendsAnomaly visibility
03Automation · enablementAI-Powered Data Engineering LabManual adaptation replaced with a two-line workflow for production-like lab testing.

Problem: Scripts required manual path, configuration, migration, and execution changes before they could run in the lab.

What I did: Built an AI agent that understands the script, adapts it for the lab, handles required changes, migrates it, and starts the workflow from a simple two-line prompt.

Technology: AI-powered engineering automation

Outcome: Reduced manual intervention and made production-like testing easier.

AI agentLab automationTwo-line prompt
04Frameworks · reuseReusable Multi-Source Pipeline FrameworkA standardized approach supporting 20+ sources without repeating the same development work.

Problem: Multiple source integrations created opportunities for duplicate development and inconsistent patterns.

What I did: Built reusable pipelines that emphasize standardization, scalability, and easier onboarding of new sources.

Technology: PySpark · Apache Spark · SQL · Parquet · ORC · Avro

Outcome: A repeatable framework supporting 20+ sources.

20+ sourcesReusable architectureStandardization
05Traceability · reliabilityBusiness Lineage & Production ReliabilityLineage, LOB visibility, and incident recovery connected technical flow to business meaning.

Problem: Business stakeholders needed to understand how an opportunity was activated or matched, while production needed recovery after a purge affected about a week of batch runs.

What I did: Innovated lineage and traceability feature tables; investigated HUB issues, unexpected behavior, activation issues, and root causes; worked with BSA and business stakeholders on resolution and recovery.

Technology: Data lineage · Production recovery · Business controls

Outcome: Clearer traceability and stabilized production operations.

Data lineageLOB / HUB issuesIncident recovery
06Knowledge · leadershipSource Deep-Dive Knowledge InitiativeA weekly Quick Espresso Call that turns source ownership and tracing into shared understanding.

Problem: Source-level knowledge silos made it harder for the team to understand the broader domain.

What I did: Initiated a weekly Quick Espresso Call where each team member owned a source, traced its flow, and presented the findings back to the team.

Technology: Source tracing · Data flow understanding · Knowledge sharing

Outcome: Better team independence and more distributed domain knowledge.

Quick Espresso CallData tracingTeam enablement

03 / profile

Make complexity
legible.

A data engineer shaped by production realities: the issue queue, the incident bridge, the source nobody documented, and the business question behind the ticket.

Working thesis

“The platform gets better when the next person can see why.”

Own the issue

Production and business issues need a clear owner, a useful signal, and a path to evidence.

Make it reusable

A framework earns its keep when it removes the same decision from tomorrow’s work.

Leave a trail

Controls and lineage make technical work explainable to the people relying on it.

Automate with care

AI-powered workflows should create leverage without moving judgment out of the loop.

04 / experience

Four environments.
One operating instinct.

01

Jun 2025–Present

Incedo Inc

Client PNC

As Retail Lead, owns a large and complex Retail segment across production reliability, LOB/HUB issue resolution, reusable engineering patterns, and reducing onshore dependency.

Retail Lead · Senior Data Engineer

02

Jan 2024–Jun 2025

Gigabyte Infocomm Pvt Ltd

Client ixigo

Designed and supported PySpark and Spark ETL workflows across AWS S3, HDFS, EMR, EC2, Hive, Sqoop, and Kafka, with responsibility for cluster monitoring, job orchestration, and streaming data delivery to S3.

Senior Data Engineer

03

Sep 2023–Jan 2024

Lumen Technologies

Client Cisco

Built Spark data-cleansing and ETL workflows, automated incremental Sqoop jobs, and applied partitioning, shuffling, and serialization practices for dependable enterprise data movement.

Data Engineer

04

Aug 2022–Sep 2023

VRJ Technologies Pvt Ltd

Client Rupeek

Developed distributed Spark applications and end-to-end pipelines using Python, DataFrames, RDDs, Spark SQL, AWS services, YARN monitoring, and Step Functions orchestration.

Data Engineer

05 / technical profile

Tools in service of a clearer system.

A production-first toolkit, with Databricks and Lakehouse engineering as the current direction.

Data Engineering

PySparkApache SparkSQLPythonHadoopHDFSYARNHiveSqoopKafkaKinesisParquetORCAvroJSONMySQLDatabricksDelta LakeLakehouse Engineering

Cloud

AWS S3AWS EMRAWS GlueAWS AthenaStep FunctionsEC2

Engineering foundation

Strong programming and problem-solving foundation. Practiced competitive coding on HackerRank · GeeksforGeeks · LeetCode.

Leadership in practice

Retail Lead, LOB/HUB issue ownership, production recovery, CR deployment ownership, knowledge sharing, and business requirement implementation.

06 / recognition

Signals from the work,
not a trophy shelf.

🏆 Polaris AwardReceived in one year.
⭐ MAD AwardReceived in one year.
🏅 Vibe AwardReceived in one year.

07 / contact

Bring the hard question.

For hiring conversations, platform problems, or a thoughtful exchange about data engineering.

The contact form will be enabled when the PHP mail handler is added.