Google Cloud Generative ai Intern
Smartinterz · Jul 2025- Aug 25
Designed and implemented a full end-to-end ETL data pipeline using AWS Glue (PySpark), S3, Glue Data Catalog,and Amazon Redshift.• Built automated schema discovery using AWS Glue Crawler and managed metadata through Glue Data Catalog forstructured data processing.• Developed PySpark transformation workflows to clean, validate, and standardize raw CSV datasets, including nullhandling, schema enforcement, and multi-format outputs (CSV/JSON).• Stored raw and processed datasets in Amazon S3 with organized folder-based data lake architecture for scalable ingestionand analytics.• Loaded transformed data into Amazon Redshift using the COPY command for downstream BI, analytical queries, andreporting use cases.• Ensured secure operation through IAM role-based access, and monitored pipeline reliability using CloudWatch logs andjob metrics.