David Griffiths is the Group Head of AI at Citi, a bank that operates in 160+ countries, serves 200 million+ customer accounts, runs a $12 billion annual tech budget, and employs 30,000+ developers.
In this episode, we discuss how he decides what to buy versus build, what it takes for an AI startup to get deployed at a bank of this size, and some of the wins from partnering with the likes of Cognition, xAI, Anthropic, and Google Gemini. We also talk about why he thinks tokenmaxxing is a really bad idea and what he is doing to keep costs under control as usage takes off.
David makes the case that there are now enough scaled use cases where smaller task-specific models will beat paying frontier prices for general intelligence, and that the timing has arrived for startups helping enterprises train or fine-tune their own models.
At our size and our scale and all the things that we do, there isn't one product to rule them all anywhere. So you have to be very thoughtful that whatever you bring into the company can solve a sufficiently scaled problem, because if it can't, you end up with this patchwork of products that you brought in, so the situation we really want to avoid is complexity.
But saying all of that, there are things that as a bank we don't want to spend time building from the ground up. It gives us zero edge and zero IP. So we would always prefer to buy something than build it, but our assessment will need to be that the product is secure and it solves a scaled problem and it's really clear to us how we would integrate it into the rest of the platform and the ecosystem that we necessarily operate.
Welcome to the Main Branch podcast. This is your host, Cagla Kaymaz. Main Branch is a show where I talk to founders and early adopters about building, deploying, and scaling world-class AI startups. Hi, David [Griffiths]. Thanks for joining me on the Main Branch podcast today.
Hello, Cagla. Great to be here.
So you're the head of AI at Citi. As we know, Citi is a global bank with 160-plus countries it operates in. It has 200 million-plus customer accounts. So this gives you tremendous responsibility and reach. Can you tell me briefly about your role and priorities as head of AI at the bank?
My role as a group head of AI at Citi, it's an interesting mandate. It's quite a new one. So we put the team together a couple of months ago. And what's interesting about the way that we're organized is we've combined the AI engineering capability, but we've also put that together within the COO organization. So we have the business mandate too to really take the capabilities that we've developed — we can talk a lot more about them in the session — but then to really land them and drive them home within the products and the business lines. So yeah, look, in terms of where we're focused, we feel pretty good about the scale of the platform that we've developed and the reach that we've got. It's with the vast majority of our employees. We get really good engagement. Most of the countries we operate within are using it. It is one platform, and we're starting to see great traction in terms of wiring this into of course our developer workflows, but then importantly into the business processes themselves.
So we're really focused on how we can become more efficient and we've done a lot of work across every business, every function, on that, making really good progress. And of course now we're thinking also about how we can put things to work more on the revenue side too. So yeah, plenty to think about.
Well, let's start with the developer workflows. You publicly talked about AI having a big impact on developer productivity at Citi. So what are some AI tools and agents that are available to developers today?
Yeah, we have a pretty comprehensive AI stack for developers at the moment. So we use a mix of things that we've developed ourselves, of course, using some of the foundational models. So we've got some pretty interesting security capabilities that we've developed in recent months. We've got a whole set of automated checks and review processes that happen across our pipelines that are triggered. So there's a small amount of stuff that we've created ourselves, which we are pretty happy with. But then also, of course, we compliment that with a pretty capable vendor stack. So we have a bunch of IDE plugins tools out there at the moment for our developers to use. I guess most notably, we've got a pretty scaled deployment of Cognition too. So Devin is out there with thousands of developers and it's been phenomenal to see that journey both in terms of how the product has evolved but also how we figured out how to get the most out of it and how our developers are now getting educated and moving beyond the early days where they were kicking the tire, seeing what the tool could do, but now really putting it to work to drive more value.
So yeah, it's a pretty broad stack. We're evaluating a few other things too, many things that you would imagine. But the thing for us, really, is anything that we use has to answer a couple of questions. One, can it scale? Two, do we have the right level of visibility around what our ... We have a very, very large cohort of developers, 30,000 or so. So it's important that we're able to see what people are doing so that we can improve the product. And so that central view into the activity is important. And then of course it's really, really critical for us that we can monitor and control the costs. It's front and center for us at the moment. That's what we look at whenever we evaluate anything to bring in. Of course, it's got to be really good, but those are the three other considerations.
Citi has something like [a] $12 billion tech budget a year, right? So I'm assuming you get pitched by vendors all the time. How do you think about buy versus build when you're looking at using Cognition versus some of the internal capabilities?
It's a great question. And I think in the era of AI, it's one of the hardest evaluations and it's been a constant set of evaluations over the last two and a half years. I think it comes down to this. Certainly if you buy, unless you are buying something that solves every single one of your problems, you're going to have to stitch it together with something else. So if you ... At our size and our scale and all the things that we do, there isn't one product to rule them all anywhere. So you have to be very thoughtful that whatever you bring into the company can solve a sufficiently scaled problem because if it can't, you end up with this patchwork of products that you've brought in, which you then do end up having to glue together with something that you've built yourself, and that's kind of the worst-case scenario.
The situation we really want to avoid is complexity because we have a patchwork of technologies and providers that we've had to somehow stitch together ourselves because of course then the net result of that for a developer or any other user is actually not a great one. Where do they go and how do you monitor things operationally and vendor contract management? So you get this massive complexity explosion if you're not really careful and selective about what you do buy. But saying all of that, there are things that as a bank we don't want to spend time building from the ground up. It gives us zero edge and zero IP. So we would always prefer to buy something than build it, but our assessment will need to be that the product is secure and it solves a scaled problem and it's really clear to us how we would integrate it into the rest of the platform and the ecosystem that we necessarily operate.
Can you make that a little bit more concrete, like an example of a platform play versus a point solution and how you would go about deciding? I'm sure there's also complexity with you have these legacy vendors that are also pitching you their AI-enabled solutions. How do you think about what's already at Citi, AI native vendors you're looking to bring in, and they might look like point solutions but they might be more of a platform play. How do they integrate? Maybe we can make it a little bit more concrete in terms of maybe use an example.
If we already have a deep relationship — and we do with many, many providers — and just because a company's been around for a long time doesn't necessarily make it legacy. I mean, you can see great products from some of these folks. If something we've already papered for and it's cost effective and we can activate something that's additive to things that we do already, of course we will do that. We won't resist it. But what we do, again, we always look at everything through the assessment of complexity. If we are going to enable a capability within a specific surface, well, is that then like a siloed experience? Does that then not know about information elsewhere in the company and therefore is a user going to have to hop between experiences to do the thing that they wanted to do? And that's always what we really want to push back on.
So those are the kinds of things we will look at. But no, ultimately if it's, well, actually we can enable some inherent AI-powered capability which is just going to improve an existing experience, of course we would certainly do that.
So we talked about software engineering already. What's another area you've seen massive gains? And to make this more interesting, let's say not customer service.
Yeah. So developers, customer servicing, and then we have a whole range of other things that we're looking at. And I guess the most interesting thing there is just the breadth of it, but it's across all of our functions, all of our operations teams. So we've got, as you would imagine, a ton of work within our businesses around how we have better document processing, how we support KYC, how we do fraud, how we do various controls, how we think about our advisory functions. And so I think the most interesting thing at the moment is it's not one or two killer examples. It's now really we are seeing dozens of examples permeate up through our businesses and functions. And we've got a very, very structured approach to that. We're literally looking at where the biggest set of expenses — concentrations of expense we have in the company — are, and it's dozens of processes.
We are going through — and we've been doing this for a little while now — going through one by one and ultimately plugging our AI toolkit in and showing actually some pretty significant results. So what, again, I'm most excited by is not one or two killer use cases. It's now we're having a scaled impact across the breadth of the company.
That's awesome. You mentioned cost cutting and efficiencies and so forth. Value of AI output can be a little bit hard to measure. How are you thinking about measuring ROI? I asked you this when we met a couple days ago, but what's your take on token maxing and measuring employee productivity through how much tokens they're consuming?
Yeah. And I'm glad that the industry has come to this consensus that token maxing is a really dumb idea, frankly. So you have to care about how much money you're spending. You do, because any company of any scale, the AI investment is de minimis. It needs its own ... It's more than de minimis, rather. It needs its own budget line. And we think about it very much as an investment that we're attaching and associating with returns. Now some of those returns are more speculative and forward-looking, but you still have to have a view on what you're going to get for the money that you're putting in. It's of course where you've got things that are just much easier to instrument and measure, it's much easier to show that you are getting a return. And I think we speak about developers and call centers, although even measuring the impact in developer land is still an open discussion in many ways.
But we kind of know the kinds of things to measure, certainly for customer servicing and otherwise, but it gets more diffuse. And I think the conclusion that we've arrived at is certainly where you are going process by process, you need process-specific metrics and measures and that's what we're doing. So you have a very specific set of baselines before and after and unanticipated effects, which you can see from early testing and whatnot. And that seems to be working quite well. Of course, where it's harder is where you have these more general-purpose tools. So our products like Stylus that we have, it's a very general-purpose AI tool, which you can do a lot with, it's great, people do incredible things with it. But associating the value of an activity in Stylus with the amount of money that you spent on the input is quite hard.
And you can see, okay, well, this person's produced this many artifacts, but how do you assess the value of those artifacts unless they're associated with a direct bottom-line-impacting activity? And you can't really see that. They're not connected things all the time, particularly in a large company. So we're thinking through how to do all of that at the moment. There's a lot of talk in the industry about outcome-based pricing, which I think is great if you can always demonstrate input to output. And that's not always possible in complex and extended processes. So I still think as a company and as an industry, there's more work and more thinking to be done on this one.
When you think about ROI for these initiatives, especially when you're bringing in a vendor, how much of that work do you do up front through a POC versus you onboard them and then over the course of a certain period, say six months, 12 months, you evaluate them and then you rip them out if it doesn't work? How much of that work needs to be done up front?
It's a great question. So we certainly have to have a view as to which part of our business will benefit from whatever it is we're going to bring in. That typically will be done then at the POC phase. And we'll try and do that very quickly. Do we feel good about the product? Do we understand it? Do we know how we would get it integrated into our processes? Are we very confident that it meets our data privacy requirements? All those sorts of things. So we do that quite quickly. But then if we could really take a position on the ROI, we will want to run it for a little bit of time, which is why whenever we would go into the deeper discussions around contracts and contract sizes, we will want to have experienced that product in real usage for some period of time.
So it's kind of a two-faced thing. Assess the fit. Does it feel right, look good? Is there a spot for it? Run it for a while. Then the proof is in the pudding there. And then we will know what the appropriate amount to invest is because we'll be more confident about the returns that we'd see.
So you're publicly working with multiple model providers like Google Gemini, Anthropic, XAI. Can you explain your strategy around using different model providers and more practically speaking, how do you choose which models to use for which use cases?
It's important that we do work with a variety of providers. The models are very good at different things, and we see that. We have a very comprehensive set of evaluations that we run internally. So certainly as the cost scenarios evolve, it's great to be able to have an option to contain or control a cost through moving it between models where that makes sense. And I think then of course for us, there's this purely practical element of just not wanting to be overly concentrated in one place. And I think as a large systemically important institution, it's an expectation that's put upon us in many ways, and it's appropriate. So working with a small and very considered set of providers has always been our strategy. It's no secret. And then ultimately, as I say, in terms of your question on model selection, it changes over time.
We're definitely seeing, as I say, the benefits of having evals that you can rerun with various models so you know when a model is good enough for a particular task or to plug in to a particular agent, for example. You can only know a model is good enough if you've evaluated it in terms of quality, correctness, performance, and cost. You just have to be disciplined about the evals that you use. It's more than just gut instinct.
There are a lot of new labs that have raised a lot of funding over the last, say, 12 to 18 months, and they have really strong teams. And some of them, their models do quite well in benchmarks. Obviously benchmarks don't always translate to practical gains, but most of the model providers you work with, these are major companies like we talked about: Anthropic, XAI, Google. How do you think about new labs? And do you think they have a decent chance of making it to the top and being used by an institution like yours?
I think the simple answer is yes, but as with anything, we need a reason to use something. What is different about the new thing? What does it displace? Does it displace a capability because it's more cost-effective? Because it's faster? For whatever the reason is. So there's absolutely space for this. And I think competition is incredibly healthy in this area. The things that we really do care about are, one, can we have a deep and high quality relationship with the provider themselves so we can learn how to get the best out of the models? Do we see that there's a lot of activity and evolution and that we've got confidence that whatever capability that we engage with will get better over time? So those are all the kinds of things that we would look at, but the bar is very high now. If you're going to go in and displace one of the frontier labs, you're going to have to have an incredibly interesting and original story, but we would be very, very willing to hear it.
What do you think about the FDE model? Does it change the calculation if they're willing to put some folks on your account and you can work really closely with them?
We want business value that comes from having the right solution engineered for the problem at hand. But I think where you have very solution-orientated FDEs who know the domain and are creative and happy to work with the technology and within the environment that we've got, I think it can be incredibly productive. But there's a huge competition for FDEs at the moment. Everyone's hiring them. Who's got the best ones will be an interesting thing to see. But no, there's definitely a space for it. And I can understand why all the labs are pushing on this now. And indeed, even ourselves, we have the same need.
We have a platform capability that we've developed internally. A big part of my organization and role is taking that and driving that into the businesses, into their processes. So we kind of have to also have FDEs to take what we've developed and land it. So yeah, I mean ultimately, if you boil it all down, you need people who understand the tools who can go and apply them. And there's a shortage of those in the world right now. The capability itself and the intelligence is outstripping our ability to deploy it as an industry. And this is why this conversation is timely.
A couple of weeks ago, we had this ... Anthropic suspended Fable for some time. I'm wondering if that changed how you think about vendor reliance and sovereign AI.
It's interesting to see how the dynamic is evolving. I think it's been clear that there was always the risk of this happening over the past couple of years. For us as a very global institution, we've had to be thoughtful about this, and we are glad that we were from the outset. And I think where we have ... Our architecture and our platform supports the ability to run differentiated models, should we need to, to give similar experiences. So that's how we're thinking about it. But of course, the more that you have specific geographical requirements around running models, then it does create complexity and new things that you have to deploy and manage. But look, we'll take it in our stride as we go forward. It's going to be very interesting to see how it evolves. But yeah, look, I think ultimately it's an architectural solution that thankfully we did the hard yards on already, so we feel okay about the future, but it's going to be very interesting to see how it evolves.
You're in a highly regulated industry so there's the vendor risk, but I'm assuming there are other things you have to think about that some of the other Fortune 500s don't need to think about. So what's keeping you up at night? What are you weighing risk-reward against these days?
I think many of the things that we do worry about I'm sure are the same things that everybody else worries about.
But of course those are concerns that I'm sure everybody has. You said this already, but in a world of agents and activity that could be to some degree autonomous, rightly we are always thinking about the risks associated with that and how we describe them and how we control them. So we spend a lot of time — almost certainly more than any other industry — thinking about those specific risks. And we have to have a lot of engagement with our regulatory community on that. So there's a lot of dialogue in the industry around this point right now. It's one of the reasons why we've embarked on things like Arc, our agent runtime, so that we've got those concerns baked in at the ground up so that we can be confident, as we do launch agents, we are satisfied that they're appropriately controlled. So I would say that's the other thing we're spending the most time on right now.
So you mentioned frontier models, the cost of running them keeps going up. What do you think of small language models like SLMs or open-weight models? In the past, I think the gap between SLMs and LLMs were too wide for them to be practical, but it sounds like maybe we have crossed that threshold. What's your take on SLMs and open-weight ones?
They have their place for sure. Ultimately, an SLM, small language model, it is less raw intelligence, it's more locked in to a specific task, but then it's going to be faster, probably going to be cheaper. So for sure as we go through and learn about where we are getting true value from the deployment of large language models and GenAI in general, as you really see a repeated problem and you can be very confident in the value that you're getting, well then of course you go in and you optimize and that's absolutely where SLMs come in. So again, as we are seeing much more maturity in the industry as a whole in terms of the deployment of AI, we now are identifying those areas that are ripe for optimization, and SLMs can be a great tool for that. So for sure, no, they have their place and I think probably more so now than 18 months ago, but yeah, they need to be part of the toolkit.
And you have a lot of data, proprietary data, that can be helpful if you were to train models. I don't know how much of this you can talk about publicly, but to the extent you can share, how do you think about utilizing your internal Citi data to help you advance?
We have a tremendous opportunity there and a unique edge, frankly, simply because of the breadth of our geographies and the breadth of the products across the franchise. So yeah, we have a few things that we're working on, and yeah, [laugh] not too much I can say on that one, but I would completely agree with you. Data is of huge value if you can connect it, if you can relate it, and if you can put it to work in highly capable AI systems. So yeah, we're on it.
So there has been startups over the last couple of years that are saying, hey, we'll help enterprises train their own models or distill, fine-tune, whatever it takes, we'll bring the expertise. And I felt like the timing was too early for it. Cost wasn't really top of mind until, say, six months ago, so people would just use the most capable model. Do you think these startups that are focused on more training use cases for enterprises, do you think their time has come? Would you consider working with one of them or do you think you have enough capability in-house for you to just run with it?
I think absolutely, I think you are completely right. It's very much a timing issue and the time, I think, is now for that. And that is a great example of something that we would look to partner on. We would see that as an accelerant. I think we have a clear idea on what we want to do and where we would deploy it and the opportunity. So for us, then, the question is, well, how do we get there as quickly as possible? And any partnerships that can help accelerate that journey, they are highly valued. So yeah, I think I would agree with you on that one.
Well, if you're advising an AI founder, whether it's where they're working on this area we just talked about or customer service or software development, what advice would you give them if they're looking to work with Citi or one of your peer institutions?
Partnership dialogue at the platform level, we don't need to be pitched. We get it. It's really like, what is the product? How is it going to operate? How's it going to work? How are we going to scale it? And very importantly, how do we fit it in with the rest of the Citi platform? How does it compose and integrate in? It's such an important point for us. If we're going to buy any AI products and platform, well, we all know the value you're going to get out of that is going to be a function of how much information and tools and integration you can perform. So that's got to be really easy or there's got to be a though process around it. So the integration story, I would say, is key.
They wouldn't know looking externally in what that would look like. That's something they would need to work with you on, right?
Yeah, absolutely. But look, if there's a product which is okay, it's a black box, take it, run it, it's a SaaS product, here's the UI, you log in, and that's that? It's just not that useful for us. We need to know, okay, well, are there bits of it we can take? Are you API first in this really core piece of functionality and capability, how would it integrate with things like Arc? It's that sort of dialogue that we would need and an openness to that — which is not always there, by the way. Sometimes it can very much be, well, this is the product, we'll try and figure out how to open up some APIs. That typically doesn't progress very well, just takes too long. So products that are designed with integration-first mindset, of course knowing that the details of that would be something that we would work on together, but that intention about the product is important for us.
Citi has done so much in terms of AI adoption and scale. What are you personally most excited about in the next year in terms of what kind of initiatives you want to go after or the capability improvements, just anything in general you're excited about?
And we're just starting to see that and we're starting to see that excitement grow in every part of the business and the belief grow in every part of the business. So as we put these platforms together, as we do things like Arc, what I'm most excited by is that we're going to have hundreds and thousands of agents doing meaningful work across the company. And it's the scaled impact of that, for us, it's just going to be phenomenal. And we can see that path ahead now. That's the incredibly exciting thing about this. It's not, well, what is this, what does it look like at the end? We can see the steps, so we just need to take those steps now, and that's a really exciting place to be.
Sounds like it's time to execute. That's super exciting.
That's right. Exactly.
Well, David, I think we're at time. Thank you so much for joining me. This was really insightful.
Thank you, Cagla. My pleasure.