Sai R. Email & Phone Number
Who is Sai R.? Overview
A concise factual answer block for searchers comparing this professional profile.
Sai R. is listed as Senior GCP Data Engineer at CorEvitas, LLC, a with 176 employees, based in Dallas-Fort Worth Metroplex, United States. AeroLeads shows a matched LinkedIn profile for Sai R..
Sai R. previously worked as GCP Data Engineer at Blue Cross Blue Shield Association and Data Engineer at Conduent. Sai R. holds Master Of Science - Ms, Information Technology from University Of Denver.
Email format at CorEvitas, LLC
This section adds company-level context without repeating Sai R.'s masked contact details.
Review company-level records connected to Sai R. before choosing the right outreach path.
About Sai R.
• Over 6+ years of experience in designing and building scalable data pipelines to collect, parse, clean, and transform data from multiple source systems and generate high-quality data sets for advanced analytics, dashboards, alerts, and visualizations.• Experience with Big Data/ Hadoop Ecosystem: Spark, Hadoop, Hive, Airflow, Sqoop, Kafka, Oozie, Databricks.• In-depth understanding of Spark Architecture and performed several batch and real-time data stream operations using Spark (Core, SQL, Structured Streaming).• Hands-on experience in GCP, Big Query, and GCS bucket experienced in handling large datasets using Spark in-memory capabilities, Partitions, Broadcast variables, Accumulators, Effective and efficient joins.• Experience using different data engineering frameworks on Cloudera, AWS, and Azure.• Performed Hive operations on large datasets with proficiency in writing HiveQL queries using transactional and performance-efficient concepts: Partitioning, Bucketing, Windowing, etc.• Good understanding of storage technologies like HDFS (Hadoop Distributed File System), AWS S3, and Azure ADLS.• Good experience in using different file formats like Parquet, Avro, and ORC in different parts of the data lake.• Hands-on work experience in writing applications on NoSQL database -HBase.• Extensive experience working with Spark, Hive, Python, Azure, and AWS suites to create Data Pipelines.• Extensively worked on Spark performance tuning.• Comprehensive experience in designing and implementing Data Lake, Data Modelling, and Data Warehousing solutions using traditional and modern data platforms.• Understanding and experience with Extract, Transform, Load (ETL) methodologies, integrating with Big Data Systems like Hadoop.• A thorough cloud professional to implement any data-related transformations using AWS and Azure cloud offerings.• Experience in importing and exporting data using Sqoop to HDFS from Relational Database Systems.• Gained knowledge and expertise in PostgreSQL which is an open-source object-relational database system, used as the primary data store or data warehouse for many web, mobile, geospatial, and analytics applications• Experienced the integration of various data sources like RDBMS, Spreadsheets, and Text files.• Expertise in developing and scheduling jobs using Airflow, Azure Data Factory, Oozie, Crontab, and Elastic search to index, fetch, and filter log data.• Good experience in Data Modelling with star schema and snowflake schema• Created facts and dimensions tables according to the data model.
Sai R.'s current company
Company context helps verify the profile and gives searchers a useful next step.
Sai R. work experience
A career timeline built from the work history available for this profile.
Gcp Data Engineer
Current• Configure, monitor, and automate Google Cloud Services as well as be involved in deploying the services using Google compute engine and Google storage buckets.• Worked in a Machine learning operations team that was building a machine learning platform to streamline the ML lifecycle including data capture, analysis, model training, evaluation, and model deployment.• Involved in designing and building modern data solutions using Google Cloud to support data visualization.• Experience in GCP DataProc, GCS, Cloud functions, BigQuery Utilizing the current state of production and determining the impact of novel implementation on existing business processes.• Developed data transition programs from Teradata to Google BigQuery using Google function by creating functions in Python for certain events based on client requirements.• Involved in building Data pipelines, end-to-end ETL, and ELT processes for Data ingestion and transformation in GCP using an App engine.• Worked on GCP services like compute engine, storage, cloud SQL, Bigtable, Dataproc, Pub/Sub, Big Query, and cloud deployment manager.• Monitoring BigQuery, DataProc, and cloud Data flow jobs with Stack driver for the environment.• Process and load Data from Google Pub/subtopic to Big Query using cloud Dataflow with Python.• Worked on developing a workflow in Oozie to automate the tasks of loading data into HDFS and pre-processing with Pig.• Involved in creating Hive tables, loading, and analyzing the data using Hive queries.• Worked on CI/CD solution, using Git, Jenkins, Docker, and Kubernetes to set up and configure Big data architecture on the GCP cloud platform.
Data Engineer
• Designed and developed a Security Framework to provide fine-grained access to objects in AWS S3 using AWS Lambda, and DynamoDB.• Set up and worked on Kerberos authentication principals to establish secure network communication on cluster and testing of HDFS, Hive, Pig, and MapReduce to access cluster for new users.• Performed to-end Architecture & implementation assessment of various AWS services like Amazon EMR, Redshift, S3• Developed spark applications in Python (PySpark) on a distributed environment to load a huge number of CSV files.• Created Databricks notebooks using SQL, Python, and automated notebooks using jobs• Implemented near real-time data pipeline using a framework based on Kafka, and Spark.• Migrated Oracle database tables data into Snowflake Cloud Data Warehouse.• Designed and implemented effective Analytics solutions and models with Snowflake.• Experienced in developing Power BI reports and dashboards from multiple data sources using data blending.• Used Spark SQL for Scala & amp, a Python interface that automatically converts RDD case classes to schema RDD.• Import the data from different sources like HDFS/HBase into Spark RDD and perform computations using PySpark to generate the output response.• Created queries using Scala, Hive, SAS, and PL/SQL to load large amounts of data from MongoDB and SQL server into HDFS to spot data trends.• Developed a reusable framework to be leveraged for future migrations that automate ETL from RDBMS systems to the Data Lake utilizing Spark Data Sources and Hive data objects.• Integrated Apache Airflow with AWS to monitor multi-stage ML workflows with the tasks running on Amazon Sage Maker.
Data Engineer
• Analyzed large amounts of data sets to determine the optimal way to aggregate and report on them using Map Reduce programs.• Responsible for data services and data movement infrastructures, worked with ETL concepts, building ETL solutions and Data modeling.• Designed, developed, implemented, and maintained solutions for using Docker, Jenkins, and Git, for microservices and continuous deployment.• Involved in loading data from rest endpoints to Kafka Producers and transferring the data to Kafka Brokers.• Worked on Snowflake Schemas and Data Warehousing and processed batch and streaming data load pipeline using Snow Pipe and Matillion from data lake Confidential AWS S3 bucket.• Involved in creating Hive tables, loading and analyzing data using Hive queries, and writing complex Hive queries to transform the data.• Experienced in design, development, Unit testing, integration, debugging and implementation, production support, AWS, Data Flow client interaction, and understanding business applications, business data flow, and data relations.• Created reusable Terraform modules for provisioning GCP resources like BigQuery, Cloud Storage, Pub/Sub, Dataflow, and more, ensuring consistency across environments.• Utilized Terraform to define cost-effective infrastructure by selecting appropriate machine types, autoscaling policies, and data retention settings.• Experience in Data Integration and Data Warehousing using various ETL tools Informatica PowerCenter, AWS Glue, SQL, GCP, Data Flow Server Integration Services (SSIS), Talend.• Developed Sqoop scripts to import export data from relational sources and handled incremental loading on the customer, transaction data by date.• Designed, developed, and implemented ETL pipelines using Python API (PySpark) of Apache Spark on AWS EMR.• Implemented data ingestion and handling clusters in real-time processing using Kafka.• Used Scala to convert Hive / SQL queries into RDD transformations in Apache Spark.
Data Engineer
• Excellent SQL Server administration skills including Database Creation, Tables, Indexes, and Clusters Creation.• Evaluated the suitability of Hadoop and its ecosystem for the project and implemented/Validated various proof of concept (POC) applications to eventually adopt them to benefit from the Bigdata Hadoop initiative.• Estimated the Name Node and Data Node software and hardware requirements and planned the cluster.• Extracted the needed data from the server into HDFS and bulk-loaded the cleaned data into HBase.• Designed, implemented, and deployed within the customer’s existing Hadoop, Cassandra cluster for a series of custom parallel algorithms for various customer-defined metrics and unsupervised learning models.• Using the Spark framework Enhanced and optimized product Spark code to aggregate, group, and run data mining tasks.• Installed and configured Hadoop Map Reduce, HDFS, and developed multi-map-reduce jobs in Java and Scala for data cleaning and pre-processing.• Managed and monitored the Cloudera cluster running on Azure.• Experienced in installing, configuring, and using Hadoop Ecosystem components.• Experienced in importing and exporting data into HDFS and Hive using Sqoop.• Participated in the development and implementation of the Cloudera Hadoop environment.• Experienced in running queries using Impala and BI tools to run ad-hoc queries directly on Hadoop.• Strong understanding of RDBMS concepts as well as Data Modeling Concepts.• Experience in BI Development and Deployment of SSIS packages from MS Access, and Excel.• Experience in implementing Full, Transaction Logs and Differential backups. • Used shell commands to schedule and execute database maintenance operations, including index rebuilds, integrity checks, and partition management.• Developed and supported tools using SQL, PL/SQL, and Linux shell scripts to maintain test and development environments, apply DDL and DML, and provide LDBA support for large-scale projects.
Sai R. education
Frequently asked questions about Sai R.
Quick answers generated from the profile data available on this page.
What company does Sai R. work for?
Sai R. works for CorEvitas, LLC.
What is Sai R.'s role at CorEvitas, LLC?
Sai R. is listed as Senior GCP Data Engineer at CorEvitas, LLC.
Where is Sai R. based?
Sai R. is based in Dallas-Fort Worth Metroplex, United States while working with CorEvitas, LLC.
What companies has Sai R. worked for?
Sai R. has worked for Corevitas, Llc, Blue Cross Blue Shield Association, Conduent, Talen Energy, and Ndiz Solutions.
How can I contact Sai R.?
You can use AeroLeads to view verified contact signals for Sai R. at CorEvitas, LLC, including work email, phone, and LinkedIn data when available.
What schools did Sai R. attend?
Sai R. holds Master Of Science - Ms, Information Technology from University Of Denver.
Search by job title, company, industry, location, and seniority. Export verified B2B contact data when you need it.
Start free trialCheck these profiles if this is not the Sai R. you were looking for.
View similar profiles