Description
Summary:
Seeking a Data Solution Architect, AWS, to lead data pipeline design and implementation for cloud-native analytics and GenAI enablement, owning architecture decisions across data ingestion, transformation, storage, and search/retrieval using AWS services.
Highlights:
1. Lead data pipeline design for cloud-native analytics and GenAI-enablement
2. Own architecture decisions across data ingestion, transformation, and storage
3. Design and architect scalable data pipelines using AWS Glue, Lambda, and S3
We are seeking a **Data Solution Architect, AWS,** to lead data pipeline design and implementation for a cloud\-native analytics and GenAI\-enablement initiative, owning architecture decisions across data ingestion, transformation, storage, and search/retrieval infrastructure using AWS services.
**Responsibilities**
* Design and architect scalable data pipelines using AWS Glue for ETL orchestration, Lambda for event\-driven processing, and S3 for data lake storage
* Define data architecture patterns aligned with AWS Well\-Architected Framework principles for reliability, performance, and cost optimization
* Create technical specifications and architecture diagrams for data platform components
* Lead development of Python and PySpark solutions for large\-scale data processing and transformation
* Architect OpenSearch and vector database solutions to support semantic search, retrieval\-augmented generation, and AI/ML workloads
* Design data models and pipelines that enable advanced analytics including regression analysis and NLP applications
* Provide architectural guidance on GenAI strategy, advisory, and operational integration with existing data infrastructure
* Design infrastructure patterns that support machine learning model training, inference, and deployment workflows
* Advise on data preparation and feature engineering approaches for ML/AI use cases
* Establish CI/CD patterns for data pipeline deployment and infrastructure\-as\-code practices
* Collaborate with engineering teams to ensure data platform components meet operational excellence standards
* Support knowledge transfer and technical documentation for sustained delivery
**Requirements**
* 7\+ years of experience in data analytics engineering
* High proficiency in PySpark
* Advanced expertise in AWS Glue, AWS Lambda, and Amazon S3
* Proven skills in Amazon OpenSearch
* English proficiency at B2 level or higher
**Nice to have**
* Knowledge of machine learning
* Familiarity with CI/CD
* Understanding of vector databases
**We offer**
* International projects with top brands
* Work with global teams of highly skilled, diverse peers
* Healthcare benefits
* Employee financial programs
* Paid time off and sick leave
* Upskilling, reskilling and certification courses
* Unlimited access to the LinkedIn Learning library and 22,000\+ courses
* Global career opportunities
* Volunteer and community involvement opportunities
* EPAM Employee Groups
* Award\-winning culture recognized by Glassdoor, Newsweek and LinkedIn
*EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.*