Senior Machine Learning Engineer, Model Risk Management
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
Technologies
Python · NumPy · Pandas · scikit-learn · LightGBM · XGBoost · PyTorch · MLflow · Databricks · Prefect · GCP Vertex AI · Snowflake
About Block (Square)
Multi-ecosystem fintech: Square (SMB commerce + payments), Cash App (consumer banking), Afterpay (BNPL), Tidal (music), TBD/Spiral (Bitcoin infra).
Public
Job description
The full responsibilities and requirements are on the employer’s site.
Open application page ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
independent function that decides whether a model is sound enough to put in front of customers and regulators. The failures that matter rarely announce themselves: a model can clear every headline metric and still be broken underneath. It can pass clean at launch and then quietly drift as the population shifts, until the loss it was supposed to prevent surfaces months later. The hard part is finding what looks right and is wrong, then proving it well enough to hold up under questioning. Much of the work arrives under-specified, so you scope it into a defensible plan, ask the questions that surface the real requirements, and defend your tradeoffs to the people who built the model you are challenging. The same scrutiny you apply to models applies to AI. We build the tooling that lets a lean team validate at scale, so you critically evaluate what it produces and own the evaluation that
More from the job description
Block builds simple, powerful tools that make progress towards an economy that’s truly open to all. Each of our brands unlocks different aspects of the economy for more people. Square makes commerce and financial services accessible to sellers. Cash App is the easy way to spend, send, and store money. Afterpay is transforming the way customers manage their spending over time. TIDAL is a music platform that empowers artists to thrive as entrepreneurs. Bitkey is a simple self-custody wallet built for bitcoin. Proto is a suite of bitcoin mining products and services. Together, we’re helping build a financial system that is open to everyone. Join us. The Role Block lends, moves money, and screens for financial crime at enormous scale, and one bad model can mean millions in credit losses, suspicious activity that goes unreported, or a fair lending violation. Model Risk Management is the independent function that decides whether a model is sound enough to put in front of customers and regulators. The failures that matter rarely announce themselves: a model can clear every headline metric and still be broken underneath. It can pass clean at launch and then quietly drift as the population shifts, until the loss it was supposed to prevent surfaces months later. The hard part is finding what looks right and is wrong, then proving it well enough to hold up under questioning. Much of th [... source excerpt omitted ...] r-lending teams who rely on your analysis, and the auditors and bank partners who carry it into regulatory engagements. This role is remote-friendly within approved US locations. You Will Independently challenge model owners across lending, fraud, and AML: reproduce their results, set and defend the acceptance thresholds, and own the call on whether a model is sound. Hunt the silent errors that make metrics lie, and prove them out before they reach production. Choose evaluation that holds up under real conditions: rare events, shifting populations, and drift that only shows up after launch. Work hands-on in codebases you did not write, learning the data, configs, and convent [... source excerpt omitted ...] ainability and fair-lending findings on consumer credit models back to the model and product decisions that follow. Help define how Block validates the systems at the frontier of production AI, setting standards where none exist yet. You Have A quantitative degree or equivalent experience, and senior-IC depth building or validating models in a high-stakes domain such as credit, fraud, or financial crime. Command of effective-challenge methodology: reproduction, conceptual-soundness review, benchmarking, stress testing, and outcomes analysis, with an eye for how a model holds up after launch and where its assumptions break. Deep applied ML and statistics across model families,
Employer postings · Data from · Sources