Principal Associate, Data Scientist - LLM Customization Team
AI summary of the role
This role is on the LLM Customization team, building and fine-tuning generative AI models for customer-facing applications at Capital One.
What you’ll do
- Partner with a cross-functional team of data scientists, software engineers, machine learning engineers and product managers to deliver AI powered products that change how customers interact with their money.
- Leverage a broad stack of technologies — Pytorch, AWS Ultraclusters, Hugging Face, LangChain, Lightning, VectorDBs, and more — to reveal the insights hidden within huge volumes of numeric and textual data.
- Be the expert in Natural Language Processing (NLP) to harness the power of Large Language Models (LLMs), adapt and finetune them for customer facing applications and features.
- Build machine learning and NLP models through all phases of development, from design through training, evaluation, and validation; partnering with engineering teams to operationalize them in scalable and resilient production systems that serve 80+ million customers.
What you’ll bring
- Currently has, or is in the process of obtaining a Bachelor's Degree in a quantitative field plus 5 years of experience performing data analytics, or a Master's Degree in a quantitative field plus 3 years of experience, or a PhD in a quantitative field.
- At least 3 years’ experience in Python, Scala, or R.
- At least 3 years’ experience with machine learning.
- At least 3 years’ experience with SQL.
Technologies
PyTorch · AWS · Hugging Face · LangChain · Lightning · VectorDBs · NLP · LLMs · Python · Scala · R · SQL
Source and classification
Internal deployment & tooling · Evidence for this classification:
edge of GenAI and at the center of bringing our vision for AI at Capital One to life. The work of the AI Training Team touches every aspect of the model development life cycle and our deployed models in production drive business impact with visibility from our C-Suite. Our team creates unprecedented amounts of high quality data for training and testing GenAI models; we care about how it’s created, what’s in those datasets, and the impact they have We are invested in building capabilities for evaluating and monitoring generative models; these methods must be state of the art, easy to use, and trusted by our users and contributors Horizontal capabilities enable vertical use case work; the team builds search, summarization, RAG, and agentic workflows for integration in production applications across the company We learn from our colleagues, attend conferences, publish papers, and
More from the job description
Principal Associate, Data Scientist - LLM Customization Team Data is at the center of everything we do. As a startup, we disrupted the credit card industry by individually personalizing every credit card offer using statistical modeling and the relational database, cutting edge technology in 1988! Fast-forward a few years, and this little innovation and our passion for data has skyrocketed us to a Fortune 200 company and a leader in the world of data-driven decision-making. As a Data Scientist at Capital One, you’ll be part of a team that’s leading the next wave of disruption at a whole new scale, using the latest in computing and machine learning technologies and operating across billions of customer records to unlock the big opportunities that help everyday people save money, time and agony in their financial lives. Team Description The LLM Customization team is on the cutting edge of GenAI and at the center of bringing our vision for AI at Capital One to life. The work of the AI Training Team touches every aspect of the model development life cycle and our deployed models in production drive business impact with visibility from our C-Suite. Our team creates unprecedented amounts of high quality data for training and testing GenAI models; we care about how it’s created, what’s in those datasets, and the impact they have We are invested in building capabilities for evalu [... source excerpt omitted ...] trusted by our users and contributors Horizontal capabilities enable vertical use case work; the team builds search, summarization, RAG, and agentic workflows for integration in production applications across the company We learn from our colleagues, attend conferences, publish papers, and maintain strong connections to the research community. In this role, you will: Partner with a cross-functional team of data scientists, software engineers, machine learning engineers and product managers to deliver AI powered products that change how customers interact with their money. Leverage a broad stack of technologies — Pytorch, AWS Ultraclusters, Hugging Face, LangChain, Lightning, [... source excerpt omitted ...] hin huge volumes of numeric and textual data. Be the expert in Natural Language Processing (NLP) to harness the power of Large Language Models (LLMs), adapt and finetune them for customer facing applications and features. Build machine learning and NLP models through all phases of development, from design through training, evaluation, and validation; partnering with engineering teams to operationalize them in scalable and resilient production systems that serve 80+ million customers. Flex your interpersonal skills to translate the complexity of your work into tangible business goals. The Ideal Candidate is: Customer first. You love the process of analyzing and creating, but
Employer postings · Data from · Sources