Back to main page

The Next Foundation Model Frontier: The Ledger

The people Tala serves are thin-file twice over. Many lack credit bureau records, pay stubs, and formal employment history. They also lack representation in frontier AI. The training corpora, including the benchmarks, evals, and data available for training, are all overweight in Western and up-market segments.

While this blind spot is true of frontier LLMs from the big labs, it is equally true for a new frontier AI that we are focused on: transaction foundation models (TFM). These are deep learning neural networks that learn to represent customer behavior from big datasets of semi-structured financial transactions: a credit, a debit, payee details, a merchant, location data, a timestamp. For these models to qualify as foundational, they have to clear a high bar:

1. Scale. Foundation models train on datasets far larger than classic ML models use.

2. Scaling-law behavior. Foundation models keep improving as you add more data and more parameters: emergent scaling-law behavior, not just marginal gains.

3. Real downstream value. Foundation models translate into real value on standard ML tasks, including fraud detection, funnel conversion, recommendations, and more, through generative prediction and transfer learning. In fintech, that value shows up directly as better approval odds and better offers for customers.

Tala serves the world’s underbanked at scale. That gives us something almost nobody else building in this space has: deep and diverse coverage among the customers frontier AI has failed to represent. We have a direct incentive to get representation right because it means better representation for the global majority and better business for Tala and our customers. 

But doing this will take a different approach than the one being taken by today’s TFMs. Most importantly, we need to expand our definition of what counts as a line item. Tala has proven for 10+ years over billions of data points that the context around a consumer informs how that person will engage with a financial product. As an industry, our foundation models need a ledger that includes transactions and the context. Additionally, as we build a foundation and framework, it needs to be measurable and serve all customers. 

In this article we’ll define the three pillars of Tala’s approach: the ledger, why benchmarks are needed and how we’ll get there, and how public and synthetic data based in reality will set the bar for this new suite of foundation models. 

Why the Ledger, Not the Transaction

Tala’s research has already shown that non-financial events can add meaningful value to underwriting and fraud detection. That insight is at the heart of what we call Ledger Foundation Modeling (LFM). The term “ledger” reflects a broader view of what belongs in a model’s training data, such as a loan disbursement, default, or repayment, but also engagement patterns within the app or a customer service reach-out and how it went. Unlike “transaction,” “ledger” makes room for the full range of signals Tala has proven useful for risk splitting over millions of data sets and loans. 

Tala uses the full ledger of events to define our interactions with customers—and our research has already shown that non-financial events add real value to underwriting and fraud detection. Our current LFM work is built to scale that finding.

Behavioral Ledgers Change Product Reads – Not Just Outcomes

Once you treat the ledger as a behavioral record, not just a financial one, underwriting and funnel measurement become two applications of the same underlying data. A loan default and an abandoned application are both sequences of behavior with context. The same representation that helps predict one can help predict the other.

That creates a faster path to proving LFM works. Credit outcomes can take months to resolve. Funnel outcomes, on the other hand, happen in minutes. We can test whether behavioral ledger representations improve onboarding or application completion before lending outcomes are visible. It makes the thesis bigger than fintech. If LFM can improve conversion, the same underlying representation has value for digital products and growth teams well beyond lending.

To Build a Rich Ledger, You Need the Data

Every new foundation-model domain starts the same way: with a corpus of information. Language models had the open web. Vision models had ImageNet. Protein models had the Protein Data Bank. In each case, a very large, very public, body of data already existed and could be used to train increasingly capable models.

Financial data is different. The most valuable financial corpus is fragmented, proprietary, and largely inaccessible. It’s sitting inside the world’s banks, card networks, lenders, and fintechs: every loan, swipe, transfer, repayment, and default. And unlike the open web, it is private by design. Banking secrecy laws, data-protection regulations, and card-network rules exist specifically to protect consumers and competitive edges. 

Excellent progress is being made in transaction foundation modeling by key banks and fintechs, but we can’t directly compare results or performance. Knowing who has the best model architecture and where to push the frontier of innovation means developing and adopting universal public benchmarks. Tala’s billions of loan behaviours, adjacent transactions, behavioural responses, and engagement data lay the groundwork for our LFM. Partner and geographic expansion lays the required data squarely in our nearterm roadmap. 

Private companies are already training foundation models on massive transaction data sets, bringing improvements to their internal prediction tasks such as underwriting and fraud detection. But the industry has no common way to compare these models or determine which approaches are truly state-of-the-art because no public evaluation data set exists. Moving the field forward requires universal benchmarks that make progress measurable and give researchers a shared view of what remains unsolved. 

The Largest Public Corpus Is the Blockchain

Public blockchains give us something financial AI has been missing: real, open ledger data. Unlike traditional financial data, blockchain transactions are permissionless and available at scale, giving researchers a place to develop and test representations on real financial behavior. Tala, in partnership with our key research partner Embed—an AI research company focused on blockchain behavior modelling since 2022—is developing a foundation model trained on public blockchain transaction data with the goal of enabling zero-shot transfer onto Tala’s own ledgers. 

Public data sets, dominated by blockchain, are the tip of the iceberg.

The question we are testing is how effectively onchain behavior transfers to consumer lending. Crypto and DeFi have their own incentive structures, their own noise—bots, wash trading, MEV—and their own approaches to identity. Whether a model trained on that distribution can generalize to consumer lending and spending behavior is a real empirical question, not an assumption.

But answering that question rigorously is itself frontier research. Cross-domain transfer requires a representation that can encode a transaction, merchant, or wallet in a way that captures the underlying behavior regardless of whether it originated at a U.S. bank or on an Ethereum smart contract. And evaluating that transfer requires the universal benchmarks described above. That combination—an open corpus, a common representation, and a universal benchmark—creates an opportunity to advance the field in ways the industry has not yet been able to. 

Blockchain data is the largest public corpus of financial transactions. Private banks and lenders in aggregate have over an order of magnitude more proprietary data over the same period of time. 

The Unlock: Synthetic, Anonymized Ledger Data

The single most promising path to open ledger foundation model benchmarks—and to meaningfully augment the training data available to everyone building in this space—is synthetic, carefully anonymized ledger data. This shouldn’t be viewed as a workaround for privacy, but rather as the actual engineering discipline that makes open benchmarking possible.

The challenge is preserving what makes a ledger useful. A synthetic dataset can match the right averages and correlations, and even produce a strong classifier, while still missing the underlying behavioral structure: bursts of activity that precede a default, the shared infrastructure that links a fraud ring, the arc of an entire repayment history rather than any single row in isolation. Most of today’s synthetic-data tooling checks whether a row looks plausible. For ledger data, the real test is whole-ledger fidelity: whether the sequences, recurrence, and relationships between entities over time hold up statistically.

Row-level plausibility versus whole-ledger fidelity: that distinction is where the real intellectual property in this field will get built.

That distinction is where much of the intellectual property in this field will get built. It’s a harder problem than most published work has treated it as, but one that is genuinely tractable with the right combination of statistical rigor and machine learning. If blockchain gives us an open corpus, synthetic data and evals give us the shared measurement approach. 

Tala’s Thesis

Taken together, the opportunity is clear. The ledger needs to capture behavior and context, not just transactions. Most of the world’s financial data is private; and the public data that does exist is too limited to support meaningful comparison. To drive the field closer to NLP’s mature state, we need three things: 

1. Ledger data that isn’t just transactions. Use all the data, not just obvious transaction artifacts, to represent financial capacity. A useful ledger should capture the full sequence of financial and behavioral events that shape an outcome. 

2. Universal, public benchmarks and evals. The field needs shared evaluation tasks for default prediction, fraud detection, contribution margin, lifetime value, and more, built on data realistic enough to matter: real transaction timing, real amounts, real context. No single company should own this benchmark. We plan to put forward a first draft, The Ledger Commons, for two reasons: to get the conversation going, and to ensure that customers like Tala’s that have been historically left out of financial AI research are represented in the evals. 

3. Representations that generalize out of domain. Building a shared benchmark only works if models can learn representations that transfer across institutions and data environments. That means developing ways to represent a transaction, merchant, wallet, or behavioral sequence consistently, whether it comes from a U.S. bank, an African mobile lender, or an Ethereum smart contract.

Tala is working on all three right now: building an eval platform, developing a synthetic dataset encompassing the ledger, and training our own internal LFMs to prove this new value. These models will continuously improve in the next few years as we add hundreds of billions of tokens from geographic, partner, and synthetic data points. The goal is bigger than improving Tala’s own models. It is to create a way for the industry to measure how well financial AI actually works for the financially underserved global majority. 

Share this article now: