About me

Hello, I'm Mohit Patil. I graduated in Masters of Science in Computer Science at Arizona State University. Dedicated Data Engineer with a distinguished 4-year track record in architecting and refining data pipelines. I am an aficionado of the intricate interplay between data and insights, transforming raw information into strategic assets. My professional odyssey is marked by a profound love for decoding the narratives concealed within the complex tapestry of data.

In the sphere of data engineering, I am an adept conductor, orchestrating harmonious symphonies of information. With a meticulous eye for detail and a proficiency in problem-solving, I navigate the intricacies of data engineering with finesse, ensuring the seamless transmission of information from source to destination.

Beyond the realm of pipelines, I am captivated by the artistry of machine learning. Delving into predictive analytics, I unravel patterns that herald the future. My passion extends beyond the algorithms, embracing the innate curiosity that fuels innovation and the thrill of venturing into unexplored data territories.

Embark with me on this expedition, where data transcends mere numbers and becomes a canvas for profound insights. Let's decode the language of data, demystify its intricacies, and pave the way for a future where information isn't just processed—it's comprehended. Join me in shaping a narrative where data isn't just a commodity; it's a strategic ally in the pursuit of excellence.

What i'm doing

  • Web development icon

    Software Development

    Professionally, I excel in designing and developing scalable, efficient, and distributed systems using the latest technological advancements.

    My passion lies in innovating robust data pipeline solutions through the exploration of cutting-edge methodologies.

  • design icon

    Data Science & Analytics

    In my professional role, I utilize data analytics, machine learning modeling, and statistical modeling to extract deep insights from complex data sets.

    My work epitomizes the transformation of raw data into actionable intelligence, driving organizational progress and achieving key goals.

Resume  

Experience

  1. Nordstrom, Inc. Hybrid

    Data Engineer II Sept, 2024 - Present
  2. Carvana LLC Hybrid

    Data Engineer, Finance Apr, 2024 - Sept 2024
    • Spearheaded the development of an ETL pipeline using PySpark and pandas to automate SEC financial reporting, realizing $100K in annual cost savings through insourcing of a previously outsourced process.

    • Optimized data processing efficiency by 90% through the migration of stored procedures from SSIS jobs to the Snowflake.

    • Engineered Tableau dashboards to provide real-time visibility into Snowflake task monitoring and data quality assurance.

  3. Copa Health Remote

    Data Engineer/Analyst May, 2023 - Mar, 2024
    • Engineered a robust Python ETL framework using cutting-edge libraries like boto3, psycopg2, pandas, polars, and PySpark, pioneering intricate data transformations beyond traditional ETL tools capabilities for enhanced data pipeline efficiency

    • Showcased metrics-driven success with substantial code reusability and a remarkable 50% development time reduction

    • Achieved an outstanding 85.6% accuracy in predicting no-show appointments through the adept use of machine learning algorithms and demographic insights, resulting in substantial resource allocation enhancements

    Data Engineer/Analyst (Co-Op) Aug, 2022 - May, 2023
    • Designed and executed ETL solution with Pentaho, gathering and transforming data from 17 REST APIs, and loading encrypted data into data lake tables for enhanced accessibility and analytics to upload data privacy

    • Crafted over 50 dynamic reports using AWS QuickSight and relational databases, furnishing stakeholders with actionable insights for strategic decision-making

  4. Symphony Health (An ICON Plc Company) Phoenix, AZ

    Data Engineer Intern Jun, 2022 - Aug, 2022
    • Extracted and transformed data from over 250 fields using up to 30 ETL components, successfully storing the data in OLAP databases to support critical business intelligence needs

    • Innovatively implemented Docker to containerize Informatica pipelines, resulting in a 20% reduction in deployment time and an impressive 15% enhancement in overall workflow efficiency

  5. Arizona State University Tempe, AZ, US

    Graduate Research Assistant at Center of Cybersecurity May, 2022 - Dec, 2022
    • Revamped 250+ test cases in Angr (a python library for binary analysis) using Python, Git, Docker, and WSL, resulting in a 60% reduction in runtime, while enhancing the efficiency, compatibility, maintainability, and readability of open-source code

  6. Accenture Mumbai, MH, India

    Data Engineer Sep, 2019 - Jul, 2021
    • Designed and implemented end-to-end ETL solutions for various databases, including Oracle, AWS Redshift, Apache Cassandra, and PostgreSQL, ensuring high-quality data models and performing thorough unit and integration testing

    • Utilised ETL tools to efficiently transform data for over 5 million users in data warehouses and data marts

    • Developed a customer-centric recommendation engine using machine learning algorithms, leading to a 5% increase in sales

    • Managed AWS S3 and global data lake buckets to securely store customer data for different regions globally

    • Conducted thorough data analysis on ETL data pipelines, identifying and fixing over 200 defects and providing data insights

Education

  1. Arizona State University Tempe, AZ, US

    Masters of Science in Computer Science | GPA: 3.93/4.00 Aug, 2021 - May, 2023

    Distributed Database Systems, Introduction to Deep Learning, Machine Learning Security, Data Visualization

  2. Pune University Pune, MH, India

    B.Tech in Computer Engineering | GPA: 8.86/10.0 Jul, 2015 - Jun, 2019

    Data Structures and Algorithms, Object-Oriented Programming, Data Mining, Machine Learning

My skills

Languages
  • Python icon
    Python
  • SQL icon
    SQL
  • Java icon
    Java
  • C icon
    Scala
  • Linux icon
    Linux / Shell Scripting
Cloud Platforms
  • AWS icon
    AWS Services
  • GCP icon
    Google Cloud Platform
  • Lambda icon
    AWS Lambda
  • Cloud Functions icon
    Cloud Functions
  • Dataproc icon
    Dataproc
  • AWS S3 icon
    s3
  • AWS glue icon
    Glue
  • GCS icon
    GCS
Databases
  • Big Query icon
    Big Query
  • Snowflake icon
    Snowflake
  • Redshift icon
    Redshift
  • PostgresSQL
    PostgresSQL
  • MSSQL icon
    MS SQL
  • Oracle icon
    Oracle
  • Teradata icon
    Teradata
  • Cassandra icon
    Cassandra
  • Hive icon
    Hive
Data Engineering and Libraries
  • Apache Hadoop
    Apache Hadoop
  • Apache Spark icon
    Apache Spark
  • Docker icon
    Kafka
  • Apache Spark icon
    Airflow
  • Pandas icon
    Pandas
  • Docker icon
    Sci-kit learn
  • Polars icon
    Polars
  • Terrafrom icon
    Terraform
  • Docker icon
    Docker
  • Docker icon
    Git
ETL and Reporting Tools
  • Databricks icon
    Databricks
  • NumPy icon
    Pentaho
  • Pandas icon
    Informatica PowerCenter
  • TensorFlow icon
    Talend
  • Docker icon
    AWS QuickSight
  • Tableau icon
    Tableau

Contact

Contact Form