← All recaps
Aug 2026 · Pillsbury Winthrop Shaw Pittman LLP, Palo Alto, CA

How to Build Your Data Layer for AI Transformation

Legal Tech Frontier Community Event #13 · Frontier Table 01 · Co-hosted with Oz Benamram · Sponsored by SKILLS.law and Legal Tech Frontier

On August 11, I co-hosted a roundtable with Oz Benamram at Pillsbury Winthrop Shaw Pittman LLP in Palo Alto: How to Build Your Data Layer for AI Transformation. We went in treating “legal AI transformation is actually a data transformation” as a premise to test. We came out with a working manual for building one.

If you want the conclusions, Oz has written an excellent summary of it, and his post has already blown up (see below). What I want to replay is how we got there, and most importantly, how you can run the same conversation even if you weren’t in the room.

We built the session around seven questions and put them to the room in sequence, under Chatham House Rule. Everyone speaks, everyone challenges. Where else do you get Sherwin Wu talking about this from the OpenAI side, Wei Chen, one of the inaugural FT Law 50, talking about it from the Infoblox legal department, and Julian Tsisin, who has been building legal technology at Meta longer than most of this field has existed, all in the same hour?

Add principals from Google, AMD and Qualcomm, and the partners and innovation lawyers now driving AI innovation for firms like Jones Day and Baker Botts.

I think Silicon Valley may be the only place this room assembles. And SKILLS.law and Legal Tech Frontier are among the first communities to put this topic on the table.

Very few firms are building a firm-wide data layer. Even fewer could articulate how. We did.

We could not fit everyone in, but this conversation should not belong to the room. Legal teams everywhere are being asked to build this, and most of them are starting without a map. I would rather hand ours over than keep it.

⬇️ The seven discussion questions we used are here. I would suggest every team taking AI seriously read them as an audit sheet.

The seven discussion topics used at the roundtable: what the data layer is, clean the lake or fish in it, structure as byproduct vs. tax, confidentiality and ethical walls, who owns it, firm data vs. client data, and experience and expertise data.
The seven discussion questions — use them as an audit sheet.

Ask your team and ask yourself this week: have we thought about this, are we already trying, and how far have we actually gotten?


The Real Work of Legal AI Transformation Is the Data Layer

By Oz Benamram — Chief AI Officer at Pillsbury | Author: Law, Reinvented | Founder: SKILLS.law | Insights: #NoMoFomOz · August 12, 2026

The next phase of legal AI will not be won at the model layer.

Firms and legal departments will increasingly have access to the same foundational models, applications, and agentic capabilities. Models are being commoditized. Context is not.

The more important question is what those systems can securely understand about an organization’s clients, matters, experience, judgment, and ways of working. That is where durable advantage will be built: the data layer.

On August 11, SKILLS.law and Legal Tech Frontier sponsored a closed-door roundtable in the Pillsbury Winthrop Shaw Pittman LLP Silicon Valley office that I co-hosted with Helen Fan, bringing together legal leaders, technologists, knowledge professionals, founders, and in-house counsel around a question that is rapidly defining the industry:

How do we build the data layer required for true AI transformation?

The market has largely moved past the “build versus buy” debate. Most organizations are now buying, building, piloting—or all three.

The harder question is where lasting advantage comes from.

One answer surfaced repeatedly: proprietary context.

An organization’s work product, institutional judgment, matter history, client relationships, and operating data are uniquely its own. The challenge is not acquiring AI. It is turning those assets into something AI can use—without compromising privilege, confidentiality, ethical walls, or professional judgment.

Six themes carried the room.

Panoramic view of the roundtable in session at Pillsbury's Silicon Valley office
The room, where it happened.

1. Connected context matters more than document repositories

The data layer is not a document-management-system migration, and it is not a more sophisticated search index.

It is the connected context behind legal work.

That context includes:

The critical word is connected.

A merger agreement, standing alone, tells only part of the story. An AI system may also need to know whether the deal was a distressed sale or a merger of equals, which provisions were negotiated or rejected, who made those calls, and what commercial considerations shaped the final language.

The highest-value legal AI questions are not “find me a document.”

They are closer to:

“Have we accepted this indemnity carveout before for a client like this, in a transaction of this type—and what was the reasoning?”

Answering that requires more than retrieval. It requires context, relationships, provenance, and judgment.

2. Stop waiting for a pristine data lake

Must organizations clean all their data before AI can deliver value?

No.

“Garbage in, garbage out” still applies, but data cleanup cannot become a decade-long excuse for delaying execution.

The practical move is use-case-led preparation. Instead of asking whether the organization’s data is universally clean, ask whether this data is reliable enough for this workflow, decision, and level of risk.

For lower-stakes research and exploration, existing repositories plus lightweight retrieval may be enough. Modern models are increasingly capable of navigating imperfect, inconsistently structured information.

Higher-stakes workflows demand more discipline. An AI system comparing 100 private-equity investment memoranda to identify recurring risk factors cannot safely rely on an undifferentiated pile of final_v2_revised_FINAL.docx. It needs to know which document is authoritative, whether it is current, where it came from, and how it relates to other versions.

The objective is not perfect data. It is fit-for-purpose data.

3. Make structure a byproduct of work

Lawyers will not maintain metadata out of the goodness of their hearts.

When data capture feels like an administrative tax, adoption fails.

The better path is to make structure a byproduct of the work itself:

Data hygiene becomes sustainable when it produces an immediate benefit for the people asked to support it. The goal is not to turn lawyers into data stewards. It is to design systems that create useful structure with as little added effort as possible.

4. Governance is core architecture, not an afterthought

Legal AI differs from most enterprise AI because access to the underlying information is tightly bounded by privilege, confidentiality obligations, outside-counsel guidelines, contractual restrictions, and ethical walls.

One principle follows from that:

Permissions must travel with the data.

An AI system should only retrieve, synthesize, or reason over information the specific user is authorized to access. Matter-level permissions, role-based controls, information barriers, source-level access rules, and logging are not peripheral technical requirements. They are part of the product.

This matters even more as systems move from answering questions to taking actions. Agentic systems may retrieve documents, initiate workflows, draft communications, update systems of record, or make recommendations that affect a matter. Organizations must be able to answer:

Auditability is no longer just a compliance feature. It is part of the data layer itself.

5. Ownership is an organizational challenge

The biggest obstacle to building the legal data layer is rarely the technology.

It is the operating model.

Who owns the data, and who determines whether it is authoritative? Who is responsible for permissions and maintains the taxonomies? Who decides which use cases justify the investment? No single function can answer all of those questions.

Responsibility has to be distributed clearly:

The organizations that succeed will not merely appoint a “data owner.” They will establish a working governance model in which different forms of ownership are clearly defined.

6. The client-data tension must be addressed directly

Some of the richest legal insights are embedded in client work—the very material firms have the least freedom to reuse indiscriminately.

That tension cannot be solved with technology alone. The industry needs more sophisticated approaches to consent, purpose limitation, data segmentation, anonymization, provenance, and value exchange. The binary between “share everything” and “use nothing” is becoming unworkable.

Not every valuable insight requires unrestricted access to underlying client documents. Firms can derive substantial value from responsibly structured experience data: who has handled a given type of matter, in which industries and jurisdictions, for which types of clients, under what commercial conditions, and with what staffing model, timeline, and outcome.

Handled well, that information sharpens staffing, pricing, business development, matter planning, and client delivery—without treating confidential work product as an unrestricted pool of reusable AI data.

A practical playbook

For organizations building a legal AI strategy, the path forward is clear:

The bottom line

#NoMoFomOz: Don’t confuse access to better AI with competitive advantage.

Your competitors can buy the same models you can. They cannot buy your institutional context.

The enduring advantage comes from building a trusted, living, governed layer around your own work—one that lets AI understand not just what happened, but why it happened, who was involved, what constraints applied, and what judgment shaped the outcome.

That is the real work. And it is where the advantage lives.


From “What is an AI-native law firm?” in Silicon Valley to “What is an AI-native legal team?” in New York, we keep asking the definitional questions before anyone has settled them.

Join us at Legal Tech Frontier, the frontier community for AI transformation in law.

📄 Read the original on Substack →