# Transcript

I have a very high-tech way of recording my talks, which is, for those of you who are still a little confused, my phone is recording. This is a lanyard that is my office badge. Taken off the badge, I've turned on the recording. Why does the recording help? I feed it to my models and then ask it to say two things: one, what mistakes did I make? Second, what did I say yesterday that is no longer true and now I have to go tell people that I have to correct myself? **One of the most important things that I'm about to say is, whatever I say now will expire very soon. It has a short shelf life—some things may last for a few years, many things will last only for a few months.** So, make sure: A, whatever you think you are learning, you check before using; and if you don't learn it, don't worry, you don't have anything to forget. Harmless.

Our course covers these topics. I'm not going to take a specific topic; I'm going to take a slice of it and talk about how AI is specifically impacting certain slices. What are the slices I'm going to pick? One is the purpose and problem framing. Another is verification. But there is a slice that doesn't quite exactly fit directly into this, but you will see the importance of this, and that is the communication aspect. **But all of these are focused on one theme, which is: why on earth are you even sitting in this class and learning it if AI can do all of it tomorrow?**

For many decades, my life was data visualization. I got my internship because of one data visualization. I created a company that does data visualization. I've even sold the company. The visualization that got me my internship at Lehman Brothers was a variation of this one, where I was taking a whole series of securities—currencies, commodities like this is silver, this is gold, this is the Pakistani Rupee against the dollar, this is Indian Rupee against the dollar, etc.—and created a correlation matrix. Each of these numbers—so what this is saying is there is a 62% correlation between the Swedish Krona and the Indian Rupee. This is a 58% correlation between the Indian Rupee and, I don't know what IXIC is, but some index.

And not just that, these are grouped together based on similarities. So some of these are negative. So you see that, for instance, gold, sorry, gold—yeah, this is gold, silver, and the Pakistani Rupee. These three are negatively correlated with almost every other currency or commodity or security. So is platinum. Now, why is it that platinum, gold, silver are moving together is probably not a major question. They're metals, okay, precious metals, they move together, but why is the Pakistani Rupee moving along with them? I don't know. Maybe they knew, but they gave me a job. They said, **"Look, this is a very different way of looking at things—not something that we had explored before, not something that we had seen before. Very good."**

When I started Gramener, which is the company we then sold to Straive, one of the visualizations that caught on like wildfire was a visualization of the Mahabharata. We said, let us take the book. Each of these is one chapter. Some chapters are long, some chapters are very small. The length is the number of words. Now let us see where different characters are mentioned. So where, for instance, is Yudhisthira mentioned? Almost everywhere. Right through the book, there is mention of Yudhisthira. Where is Arjuna mentioned? And we click on this and say, okay, Arjuna's mentioned in many spots, not necessarily the entire book. And that's when I realized that there are whole portions of the Mahabharata which have nothing to do with the war or the brothers but are just general philosophical discussions—Yama and Yudhisthira talking, stuff like that.

Then I had this doubt: okay, wait, there was this little episode about Karna and—where is Draupadi? Draupadi, Draupadi, Draupadi... Ha. Having an affair. Now, is that true? I scanned through the whole thing and found that there are very few chapters where they're even mentioned together. So at least according to the original epic, this can't be. **Now, this is a very efficient way. I don't have to sit and read the entire story to find out the answers to important questions like these.**

But then that got me thinking: hold on, then I can start connecting several characters, right? Because I can then ask a question: how is—who's selected? Okay, somebody else is selected. Is that Draupadi? No. Bhishma? No. Let me just reload the page. Yeah, I can start asking questions like who is Nala? I don't need to know the epic, but there is a short segment about Nala. So, fine, there is a short story about him. Let's take Damayanti. Damayanti is appearing in exactly the same spots as Nala. So Nala and Damayanti are more closely connected than, let's say, Nala and Subhadra, who's appearing in very different spots.

That means that maybe I can start creating a similarity of characters, how closely they are related, etc., and explore prominent relationships. And that's what leads to this one, where I can see that Yudhisthira is a very centrally connected character, very close to Krishna, Pandu, Arjuna, etc. Arjuna is closely connected to Vaishampayana, whereas Vyasa is a slightly more peripheral character. Gandhari is talking only to three people: Vidura and Dhritarashtra—brother-in-law and husband—and Kunti, sister-in-law. Otherwise, she's not really connected. And allowing me to answer questions like: who was Draupadi's favorite Pandava? So, here is Draupadi, and Sahadeva is close by, Nakula is close by, Bhima is close by. Yudhisthira's not too far away. Arjuna is the furthest away.

I said, "Wait, hold on. I thought Arjuna was supposed to be her favorite Pandava. This is telling me that something is wrong." I redid the calculation. Redid the calculation again. Yet again. Every way I tried it, Arjuna was the furthest away from Draupadi. Until one investment banker in Pune said, "Anand, have you heard about this Bengali play which in turn is based on a Marathi play, or the other way around, which is about how actually Draupadi likes Arjuna but there's no mention of Arjuna liking Draupadi?" And if you are trying to plot a symmetric relationship, you should realize the entire play was about how Draupadi felt ignored because Arjuna went around gallivanting with lots of other people.

So, okay, **not only am I learning about the Mahabharata, I'm learning about derivatives of Mahabharata and inner politics and so on.** And I would share this with lots of people, and people would say, "Oh, data visualization can show you so many things," etc., and started buying our services, which obviously really helped a lot. But here is the thing: **today, AI can create these visualizations faster than I can.** The bulk of what I used to do for about two decades was painstakingly create these visuals, write the code for it, iterate manually. Now AI does it significantly faster. What does that mean? Does that mean you don't have to learn data visualizations?

Okay, actually, quick show of hands. How many people think therefore we still need to learn data visualization? Okay, that's about... how many people think we don't therefore have to learn data visualization? I see three and a half... okay, four brave—any more hands? Okay, four brave hands. I have no idea who's right, who's wrong. This is one of those things where I have zero clue. But who knows, maybe the four of you are right. Maybe we don't need to learn data visualization, or at least what we think used to be termed data visualization is not.

**So the first part of my session is about why you may not need to learn data visualization.** Firstly, because AI can do it a hell of a lot faster. How exactly does it do it faster? This is—an interestingly interactive session. What I mean is, please stop me at any point. But otherwise, like Kareena Kapoor said in *Jab We Met*, I will just keep talking. [Laughter]. So please raise your hand or interrupt at any point.

The Times of India came to me with an interesting problem. They said, "We have this property called 'Stat-Toistix'—statistics with a 'TOI' instead of 'TI' in the middle." And they said, "We are going to shut down this property." Why? Because it takes a lot of effort to create these data visualizations. I said, "Okay, fine, let me do one thing. I will give you five interns. You do them." Those five interns were agents. What do I mean by agents? ChatGPT.

What did they create? If you've seen Stat-Toistix in the Times of India published, here are some of the properties that they carry. These have been 100% AI-generated. How were they generated? Step one, I put in the following prompt: "Download all data from UNdata—data.un.org—efficiently. If there's an API, you find out how to use it, write the Python script, blah, blah, blah. Make sure that the script is resumable." This is the entirety of the prompt that I first gave maybe Codex, which is a coding agent. And it downloaded all of the UN data.

Earlier, I would have had to figure out what is the API, how do I retrieve it, write a program, download it, run it, etc. My effort was approximately five minutes of looking at this, maybe a minute or two of checking if it's still working or not. And its effort was half an hour of writing the program and maybe half an hour of running the program. And the data engineering problem, which used to be a big part of data visualization, seemed to vanish. Definitely faster.

Second, after it did that, I said, "Analyze this for newsworthy insights." And I said, "Look, people would have already analyzed this a lot. What I want you to do is find obscure stuff, niche stuff, surprising stuff, things that I can publish in the Times of India, but also obviously make sure that it is true and verifiable and simple enough for a lay audience." Again, five minutes for me, one or two hours—this may have taken more than a couple of hours—and it came up with a whole series of stories which, when I shared with the editors of the Times of India, they said, **"This is better than anything our journalists produce."**

**Not only is it faster, it is also better.** And I will talk about the "better" part of it. They said, "There's only one thing: we want to check if it is correct. How do we know that it is not making a mistake? Our journalists also make mistakes, but the journalist I can ask them, 'How did you do this?', sit with them, blah, blah, blah. You have given me something, I don't know how to do that." So—sorry, before that, there was one more step, which—ha, okay, no. Yeah, the other step was: "Render them as consistent SVGs." I basically told one of these agents, "Go to the Times of India, download all the PDFs, see what they look like, and create an instruction file called 'Stat-Toistix format'." In that, you put in an explanation of how the cards that they create on paper look, and make sure that you create them in the same way, in this particular format. And it created all of these charts.

Did I have to make any edits to these charts? Some of them. In some cases, the text spilled over. But then they reached out and said, "Anand, you don't bother doing all this. Leave some work for us. We have a graphics team. That is what our graphics team does anyway, and we are very efficient at it. We will take care of all of these things; you just give it raw." "Okay, fine, why take your job?"

And they said, "We want to verify." So I told it, "Look, I want you to add citations so that Times of India can easily verify the facts. And think about how they will verify it, and give a comprehensive verification checklist and a statement of procedure." So against each one of these, there is a little "Verify" button. If you click on it, it gives them a step-by-step process by which they can verify. It says what this card is saying is such and such, these are the fields that I got it from, and here are the tables that I got it from, here is the verification of the source table, how many rows it has, how many columns it has, how many values it has. Every piece of text in the card that has a number, here is the value that I got from the source. So here are the ways in which you should verify, and here are things that you might get wrong.

Finally, here is the checklist. You go confirm if the actual number of records was actually 41,659, the government hospital net is this. As you mark each one of these, put a tick against it. If you have confirmed all of these, then the card is correct. They followed that procedure for the first card, for the second card, for as long as probably the tenth card, and then they said, **"Yeah, okay, this is not making a mistake. Maybe we will check every third or fourth; we will run a sampling of sorts." Effectively, what we were doing was compressing the cycle of identification, generation, and verification.** What is the insight? How should it be presented? And how to verify?

If that is the case, what am I doing? Writing those prompts? I've written those prompts. Now I don't even have to do anything. I just have to say, "Do it again." And that is how the rest of the cards were created, actually. Running automatically. Why do I have a job? It visualizes better than I do. One of the things that I've been struggling with for a very long time is, how do I create a visualization of all the *Calvin and Hobbes* strips?

So **my life's greatest achievement till date is that I sat and spent six and a half years typing out every single *Calvin and Hobbes* strip.** And that resides in a secret website which I don't publish because I already got a Digital Millennium Copyright Act takedown notice, but effectively I can now search on this for, "Oh, what is that little Tracer Bullet comic where, ah, yeah, one of these, maybe this one or this one, maybe the next one," and I can search. This is something that I was doing about two decades ago.

But I wanted to create a data visualization out of this and I had no idea how. I had an idea—see, some comics are similar. I want to group them together. I want to see if there are little clusters of these comics. But there are 3,000 of them. That's massive. I don't know how to do it. So I failed at that attempt. Agents are there, so I gave the same problem to Codex, and this is what it created. Each of these little dots is a *Calvin and Hobbes* strip, and grouped by similarity. So this, for instance, is a very unusual cluster—very different from all of the others. And this is where the class bully, Moe, is, well, bullying Calvin. Completely different series from the rest, partly from a style perspective, partly from a text perspective.

Here's another completely different series. These are all the super—whatever—Calvin imagines himself to be a detective, to be Spaceman Spiff, to be a dinosaur, to be Superman, whatever. All of those are a completely different cluster altogether. Another very different cluster is about the babysitter, Rosalyn. For those of you who know *Calvin and Hobbes* can relate to this; those of you who don't, I'm just having fun, okay? But the point is that **A, it is able to create a similarity matrix. B, it is able to create a visualization in a way that I would not have thought of.**

This is a UMAP, for those of you who are aware. It effectively does the equivalent of a principal component analysis. And if you are not aware of that, then it basically says, "Look, here are the two main ways in which the comics are different," some two axes that it's constructed, and it is plotting against those two axes. And it says, "If you want three axes, I can plot it in three dimensions and so on," and gives me a timeline to show how these evolved. So if I take a one-year timeline and play this, this is how the comics evolved. So I can watch this cluster, for instance, and say, "Okay, right across, it was there, but then it vanished in the middle, and then Moe came back... okay, the imaginary characters came back."

**Student**: Is it using similarity based on the text and description in the strips or is it something else?

It is based on the text and image embedding using the Gemini API. What does that mean? **Embeddings are where you convert text into numbers.** How do you convert them? One way of doing it is to take, let's say, all the 40,000 words in English, put it into Excel and say, "This particular comic had these 40 words: tick, tick, tick, cross, cross, cross. Next, it has these words," etc. So now I've converted it to a binary 40,000-digit number, right? Except that instead of English words, what if we took concepts? The trouble with words is that "cold" and "cool" are different words; there's nothing similar between them. But if I thought of it as "coolness," is there a 70% similarity to it? I'll put a 0.7. "Anger," is there a 5% similarity to it? I'll put that number.

So it creates a set of numbers, and it has a diminishing importance. So Gemini by default, I think, supplies 3,000 numbers from zero to one that represents what this is, such that if you take the similarity, the distance between any two of these pieces of text, it tells you how similar they are. And it also lets you pass images, and it tells you how similar they are, again, based on those concepts. So I took the images and the text and said, "Give me the similarity."

This, therefore, firstly is not something I would have thought of because I didn't know about embedding similarity. I would not have thought of because I didn't know about UMAPs. Also, something that I would not have thought of because I wouldn't have even known that this sort of a thing can lead to temporal patterns. All of this emerged because I just told the coding agent, "Here is my broad requirement," and it generated something that is better than what I am capable of.

Which extends to so many areas. Because this is a representation of the images of different data visualizations of publications. Each of these is one data visualization publication. These are from the *South China Morning Post*. This is one of the pieces that they published, this is another piece that they published, and you can see that they have a certain character—they're kind of similar. And their embedding similarity puts them all together, which is very different from these publications from the *Wall Street Journal*. Their charts are completely different.

And you can see at a glance that if I take the *South China Morning Post*, it's a very different cluster from, let's say, let's take, I don't know, *Reuters*, which is a very different cluster altogether. **So if I had to say, "Give me two completely different styles of data visualization across all the papers," then these two, I would say, are very contrasting things.** This is not a question that I even knew I could ask, but it is able to solve it for me without my even asking for it. And similarly, *The Guardian*'s kind of all over the place, *Financial Times* is all over the place. So I can see that they adopt different styles of visualization depending on the context, but *South China Morning Post* is truly distinctive in that regard, and so is *Reuters*.

I can use this for research publications. What we have in blue here on this UMAP is the set of—think of it as all publications on OpenAlex, all research papers. National Institute of Engineering publishes research. Those research papers are highlighted in orange. So I get a sense of where they are publishing. And now I can start advising them on, "Look, here is how the research papers have been trending." And I'm not going to go through the details of this, but—okay, this is probably too short a time period. It will eventually start—yeah, okay, no, I'm going to move this manually and say, "Okay, look, you really haven't been publishing in the early days. Okay, most of your publishing has started only recently."

And if I look at the scope of your publications... okay, this is too slow or my publication dates are wrong. But effectively we were able to tell them something along the lines of how their field is evolving, taking this. On the same UMAP, we could see that over time, mathematics had significantly evolved. **One axis, the X-axis, was roughly on the left side: how much the field is focusing on AI, vision, perception, etc., versus materials, energy, and devices. Stuff that is more to the right is more on the materials, energy, devices side.**

Y-axis is on the top, more towards living systems and biology, and at the bottom is more towards equations and formal systems. Mathematics from 2021 had moved away from living systems and biology. Basically, a lot of COVID research was on the biological side; after COVID, they started moving away and going towards more formal systems and equations. And away from AI, vision, perception, which is interesting, towards materials, energy, and devices. So if that is the case, then you as National Institute of Engineering should also start moving your focus towards where papers are getting published instead of being left behind by papers that were written two, three, four years ago. Is the kind of thing that we could draw. So it actually has real-life uses as well. Again, something that it was able to do better than I could because this is 100% AI-generated.

It also visualizes more creatively than I do. I said, "Create a visualization—create a bunch of visualizations that I have never created, nobody has ever created, genuine innovation, and show me what you can find." This is probably about four or five months ago, and here is something that it created: a causal lag clock matrix. Now, this is something that I'm particularly fond of because it is very similar to the visualization I showed you at the beginning—the scatter plot matrix which got me my first job, well, first internship job.

You have a bunch of variables: yield, I don't know what HSTS stands for, return, consumer price inflation, etc. But it takes a bunch of financial factors, eight of them, and asks which of these are correlated with which. That is a correlation matrix, we know that. But whether one is a lead or a lag indicator of the other. And that it is showing based on this little clock. So, the correlations are negative if you see it in red, correlations are positive if you see it in blue. But this clock-like thing, if it is turning towards the right, then it means that the column is a leading indicator. And how far it is to the right tells you—so for instance, here it is saying yield spread is a strongly leading indicator of retail sales by about seven months, which means that **if you want to predict retail sales seven months later, look at yield spread today.** The correlation is as much as 58%—not a bad correlation. That is useful.

So then I can start looking at where the clocks are tilted towards the right side and say, "Okay, those I can use as leading indicators." If I'm currently using something to predict another variable and I find that it is a lagging indicator, I will stop that. **This is a novel visualization in that A, it does not exist, and B, it is useful** in that, well, I actually do plan to use this wherever I can. Now, completely invented by one prompt which said, "Create something for me that nobody has ever seen, etc."

Okay, but all of this is generation. I'm sitting here teaching a class. At least as a teacher, if I'm correcting your papers, I need to learn how to critique your visualization. Or at least for that, I have to know data visualization, maybe? So one of my colleagues, Jaydev, he sent out a WhatsApp message saying, "Here is a particular chart. Guys, any critique on this?" I'm not going to go into what the chart is just yet, but I prompted Claude in this case saying, "Look, a friend is asking me for some critique on this. Give me the critique and, by the way, give it to me as HTML."

And here's what it said. "See, he has created a dual axis, and what that means is that the top chart and the bottom chart are difficult to connect. Here, the Y-axis is truncated." And then I had to peer closely. "What is truncated? Oh, it starts with a 0.06, not a zero. Correct, that can lead to misinterpretation. Instead of 0.11, why didn't you say 11%? Very true, makes it so much easier to understand. X-axis... okay, free electricity, midday meals, LPG subsidy, and you're putting a line chart? Is there some kind of order to this?"

This was probably the best point that I could have highlighted. "What is the meaning of this color? And why are you using red if it is not a bad thing? Correlation versus causation." A series of points: A, more than I could point out; B, better than I could point out; C, certainly faster than I could point out.

**Why do we need data visualization at all? Meaning, in all of these cases, we're saying AI can create better data visualization, AI can critique better data visualization. Boss, why do you need data visualization?** One of my colleagues was working at a logistics company and he used Snowflake, which has Cortex embedded inside it. Cortex is like the ChatGPT sitting inside Snowflake; Snowflake is a database. It has all of their company data. And he said, "Give me some use cases for this company." It created 15 use cases.

He was very happy. He took it to their head of analytics. "Look, I've created 15 use cases." The head looks at this, looks at another report and says, "Thanoj, we created a... sorry, we asked a strategy consulting company to come in, and in a three-month exercise which was very expensive, from November to February, they had come up with a series of use cases. There is an 80% overlap."

Okay, Thanoj got a huge kick. He said, "Very good, now I can identify use cases." So he said, "Now I will create data visualizations solving each one of these." So revenue forecasting I will do this, pricing anomaly detection I will create an outlier chart, and so on. At which point I was asking Thanoj, "Yaar, what is the guy going to do with that pricing anomaly detection?" He said, "Well, he will correct, he will find out where the anomaly is and then correct it." "You do it for him, no?"

So he sent out this email which said... where did this go? Yeah. "I ran an anomaly detection check and for our services which cost $175 to $179 every month, there were 105,000 transactions which are less than 16 cents. Either you have a massive data quality problem, or you are charging 16 cents instead of $160 or $170. In which case you have a massive revenue leakage. Either way, you have a massive problem."

Ultimately, this is what they need to know. Why do they need to see the visualization? Same thing for each of those 15 use cases. He had created a series of visualizations and I said, "Don't bother, just send them the email." And for instance, one of the emails was, **"We are losing $170 million in revenue. Here is the fix." That one subject is more important than any visualization that you can show. $170 million.** And explaining, "Look, go talk to US Department of Energy, Meridian Chemical, Summit Refining, etc. If you want to retain some of these customers, otherwise they're leaving."

That is a case. Meaning, see, what do we want people to do based on data visualization? Decide something and act on it, right? You decide for them and you tell them what to do. Better yet, you act for them. What are agents for? Connect it to whatever sources to fix the problem, or at least tell them what they need to do. Why do we need this course? Any guesses? Or any thoughts, any counter-thoughts? So here I'm basically saying, look, you don't need this firstly; you don't need to learn data visualization because AI knows data visualization. Secondly, you don't need the course at all because we don't need data visualization. Thoughts?

**Student**: What I think is basically every time we are going through something, we are visualizing, whether in text, whether in table format, or whether in pictorial form. So in that sense, now what matters is how fast are you going to convey a large amount of information in a short span of time? How better you do that is what, you know, decides how you're going to, you know, kind of harness some amount of information.

Visualization efficiently conveys understanding. Why do you need to understand? Agent is understanding. Why are you standing in the way slowing things down?

**Student**: It's always being driven by human beings.

An excellent point. **Ultimately, we don't yet have a mechanism to hold an agent responsible. Humans are responsible. Companies are responsible.** We know how to take companies to court. Even ships are responsible. Ships can be taken to court, stripped of their assets. Even rivers can be responsible. Temples can be responsible. There is a Māori river which is a legal entity. Temples have their own funds and can be administered. We've cracked all of that.

We have not yet cracked how to give agents money, how to take an agent to court, how to kill an agent, or in the case of a company, we dissolve the company, we shut down an agent, whatever. We will figure it out. So until then, humans are responsible. So maybe for a few years, maybe even for a few decades, maybe in some places, maybe in lots of places, at least if we are accountable, we need to understand data; visualizations can compress that. God. But if accountability is not required, why else do we need visualizations?

While you're thinking about that, I'll flag off that if accountability is important, meaning you're hanging your neck out for it, your need to verify stuff becomes important. Is the visualization actually showing you what you really ought to see? We'll keep that in mind. Any other reason why visualizations are important?

**Student**: To view some unstructured data?

Why do you want to view unstructured data?

**Student**: To get some insights out of it. To just get to know what exists.

Why do you need to know what exists? You tell it to get to know what it exists—what exists. No, no, you may be leading to something. Go on, think about it.

**Student**: Decision making.

Decision making, okay, but agents can make decisions.

**Student**: Or to verify like whether our decision is right. We take decision and check whether it correlates with things.

As a cross-check, got you. But if you find that the agent is taking decisions as well as you do, and these days it is smarter than us, then? Where you're stuck may actually be the point, which is, "Look, I am not sure I know, but maybe there is something." Put another way: **I don't even know the question to ask. So how do I know that it is giving the answer when I don't know the question?**

So part of a visualization is to say, "Look, I want a different perspective of looking at what I don't know." If I had a question, I would ask it, yes, without being able to visualize, or it visualizes and sees the results and gets to the answer—it will solve the problem, okay. But when I don't even know the problem? That's a fair point.

I'll tell you how I'm using visualizations or why I'm using visualizations and where I'm using visualizations. And this is part of... so, part of what I was explaining was for... okay, and part of what I am explaining now is what you need to keep in mind when you are framing a problem, or what is the purpose of data visualization in the first place. And we will soon come to how you go about verifying this and the missing communication element to this.

But **one of the reasons I use data visualization is to understand stuff.** Now I was saying, wait, why do you need to understand stuff when the agent understands stuff, right? Because there are some things that I do for myself. I eat. I sleep. The agent can't sleep on my behalf, can't eat on my behalf. The agent also can't watch movies on my behalf—well, that's coming later.

The agent can't solve my curiosity. One of the things I was curious about is... so one of my friends, Karthik, he published a blog post about Bangalore weather. And there was something about rainfall, so that got me thinking. Is there a specific time in which it rains in Bangalore? Got me further thinking: maybe there is a specific time so if I just carry my umbrella in that time, I'm good; outside of that, I don't have to carry my umbrella.

And maybe there are tighter windows for some cities, weaker windows for some cities, more chaotic, less chaotic, more predictable weather. So it'll be interesting to see if there are cities where across the year, there is a tight window where it rains and outside of that you don't even have to carry umbrellas. So I tried looking at it, it gave lots of stuff, but I found that the best way for me to understand this was this visualization where Caracas in Venezuela has the tightest window.

Across January to December, the color tells you how well you can predict the umbrella period. Meaning, how much it will rain only inside the umbrella time period and will not rain outside the umbrella time period. For instance, in August in Caracas, there is as much as a 57% chance of rain inside that window and outside that window only a 5.7% chance. That is about the best kind of segregation—just that tiny window. And I can look at the chart and say, "Ha, okay, yes, makes sense." The yellow is the umbrella window and the green is the non-umbrella window, and yes, it is clearly very tight and it is clearly true right across.

**This part is the information compression. I can instantly understand and I can instantly verify, and that is useful.** Now, is this a standard chart? Maybe, maybe not. I wasn't even thinking twice about it. I knew my problem clearly. My problem was: I want to figure out in an instant where it only rains inside that umbrella period, and I will know it when I see it. That helped.

In contrast, for instance, in Ethiopia, there are some time periods where it's... there's a strong umbrella zone and not otherwise. Let's take Kolkata, not too good. Let's take... look at it from the bottom. Which are the worst? I should just... yeah. No, I should scroll all the way down and... a city like Cairo, it's hopeless. There is no umbrella period. In Baoding in China, in most of these places... I'm just checking if there's some place in India or which is the worst in India. Ah, okay, no, not yet here. India is not too bad. China, Chicago... why am I doing this? And this is why I later on also added a search.

And Ahmedabad is one of the worst. You just cannot predict and carry an umbrella at specific times in Ahmedabad. But Pune is not too bad. At least in June, July, August, September, you can carry an umbrella in specific periods. Kolkata is okay. In Chennai, it's not very good. If you are... if it's a rainy season, you have to carry an umbrella right across.

Now here is the second part. I'm able to tell this to you as a story and you can look at it and say, "Ha, okay, makes sense, right?" And we'll come to this communication part as well. **It is not just for your understanding; it is for the other person's understanding as well.** And that other person's understanding starts becoming more important. But this is one reason: there are times when I do need to understand because I am the actor, I am the consumer, I am the endpoint. There's nothing that I need to do. Me consuming is the purpose. Fair.

I also use this to explore. I watch movies. I want to know what movie to watch next. And this is the list of movies on the Internet Movie Database. This is a visualization that dates back to 2008 when I was talking to Col Needham, who's the founder of IMDb, and he and I were getting into a discussion on how we pick movies to watch, etc. And we just sketched on the whiteboard on how we might plot on the X-axis the number of votes on a logarithmic axis that a movie has got, and on the Y-axis the rating.

So out here on the top right are movies like... okay, you can't see those. Let's zoom out a bit and... you should be able to see some of these. Let's... yeah. So Breaking... okay, you can't read it, but I'll read it out for you. **_Breaking Bad_ and _Game of Thrones_ are right up there.** But that's series. Let's take only the movies, and out here is _The Shawshank Redemption_ and _The Dark Knight_ and so on. But wait, you'll say, "Oh, that was last decade, decade before that." Fine, I'm also curious about what are the new movies. Nothing that good—still a long way to go.

Okay, maybe people just haven't watched it enough to give that many votes. But I can find movies like... let's see what we should be... okay, what's your guess? What is the movie that is on the top right?

**Student**: _Spider-Man_.

Sorry, is that what I heard? _Spider-Man: Into the Spider-Verse_? First or the second? Okay, my guess is neither because it may not have had enough time to go up to the top, but okay, _Spider-Man: Into the Spider-Verse_ is one guess. Any other guesses?

**Student**: _Dune_.

_Dune_, okay, possibly. _Avengers_ is a possibility, but _Avengers_ is 2020s or 2010s?

**Student**: 2019.

Okay, _Avengers_... not sure purely because of the timing. _Dune_ is a possibility. Any other guesses?

**Student**: _Oppenheimer_.

_Oppenheimer_, okay. Any other guesses? Chalo, let's find out if _Dune_ is there. Okay, _Spider-Man: No Way Home_—so maybe you meant _Spider-Man: No Way Home_, not _Into the Spider-Verse_. And _Oppenheimer_ is also there. Yeah, spot on, you know the movies to watch.

I have not seen _Oppenheimer_ fully yet, only halfway there. But this is how I pick movies to watch, but not just movies to watch, but also outliers. **What is this movie? Very popular, terrible rating. Almost an embarrassment. Any guesses?**

**Student**: Indian movie?

Indian movie, okay. Any other guesses?

**Student**: _Snow White_?

Any other guesses? _Snow White_ is in fact correct. [Laughter]. Now, this is not just a disaster of the 2020s, this is the disaster of all years. Like, there has never been this popular and terrible movie ever.

But here's the thing. Because of this, I said, just after almost two decades of this visualization, I said, "I'm interested in plotting the outliers. What are the movies that are sitting out there as outliers in each decade and how does this vary, and so on?" And that is a useful point of view.

But for me, again, this visualization is helpful because it actually serves a useful purpose. I need to consume the visualization, and giving it to me in a table and all will not... and there's no way I would have spotted this. And that is partly because of the exploration. There are so many things I can explore. So for instance, I can look at animation movies and see how it evolved over the years.

So in the 1930s, there was only one animation movie, _Snow White and the Seven Dwarfs_. That's it. That is when the genre started. 1940s, you can see two clusters. There are four Disney movies in that decade and four non-Disney movies in that decade; others started. And you can see the clear difference: **Disney was the only show in town as far as animation goes.**

1950s, even fewer, _Cinderella_ and a whole bunch of these, all Disney. And 1960s is when things start getting a little shakier. Again, some popular Disney movies. 1970s, however, Disney loses its charm. These are not all Disney movies, for instance, and the ratings are lower. But the revival comes with movies like _My Neighbor Totoro_ which incidentally is a Miyazaki movie—the Japanese studio started pushing up. _The Little Mermaid_ happens to be a Disney one, this is a non-Disney one.

So there is a revival but from different circles. And then there is the Pixar and Disney revolution. Disney with its last great original movie, _The Lion King_, and Pixar with its first new ultra-blockbuster, _Toy Story_, changed the game altogether. And then it moves on... just look at the sheer number of movies that started with computer graphics becoming popular.

So **effectively, I get to see the history of a field, a question that I didn't even have**, which is the kind of thing that we're also looking for. Which is, I didn't even know that I had this question or that this is a thing, and I'm able to explore it. So one, I want to understand because I am the person consuming it. Second, I want to know what I don't know and explore something completely new. That happened in... okay, I'm going to skip this in the interest of time.

There are many ways in which... am I taking this here? Actually, I'm going to take this in... yeah, okay, fine, let's talk about this. **There are two ways in which AI can create images. One is you tell it to write a program**; it will write the HTML, JavaScript, SVG, whatever. **The second is you ask it to create a raster image**; it will create a raster image. In the first case, it is using its knowledge of coding to create stuff, that means it can be interactive, editable, etc. In the second case, it uses all of the images that it's been trained on. It is somewhat differently editable—you have to tell it to make changes to it which it may or may not do well—but there is no restriction or limitation on the kind of things that it can create. Anything that can be represented as pixels, it can draw.

Now I have a problem. I don't even know what I can create, but it does a good job. Now here's the thing. Earlier I would rely on my creativity. Now, I still need to express creativity to other people because people want to see things in a creative way—that's a valid need. So how do I go about bringing in that creativity? What are the different ways, for instance, of communicating information in a simple way?

So I had Claude generate a series of prompts saying, "Show me the different ways in which you can render comics." So it said there is an elegant brush style. There is a spot black economy style of comics. There is a flat icon style of comics, a ratty line, and a whole host of these. These are the black and white ones, and it goes on to... okay, here is the _Peanuts_ style, here is the _Tintin_ style, here is the _Archie_ style, _Amar Chitra Katha_ style, blah, blah, blah. So now I have a catalog where I can just copy and paste it saying, "Generate it in this particular style."

And it can create new styles also, obviously, like you saw it creating new data visualizations. **The reason I bring it in here is when it comes to exploration, it is not just exploration of the content that we are asked to do as data visualizers. We are also asked to explore formats.** What are the different mediums that we should represent in? Why should a data visualization not be an engraving on the factory machine out there? Literally an engraving.

Why should it not be a bunch of beads that the device automatically drops as it produces, so that across five little piles, you see a little stack over time on how many errors it has made of different kinds? Physical data visualization. **Why should it not be data sonification?** Depending on the sound that it makes and the pace at which it makes the sound, you get a sense of the data. Ultimately, operators listen to the machine and get a sense of how it's functioning, right? That is a feature, not a byproduct. You can design for it. These are different kinds of formats that you can start bringing in, which takes us towards communication as being an essential reason for the existence of data visualization.

Now, by now you've gathered that when I say visualization, I'm mostly talking about visual representations, but really data representations of any kind also I broadly club under data visualization. Peripheral field, borderline, not something that we are covering in this course, but just something to keep in mind. Which brings us to the communication part. **If I have to explain stuff to people, then clearly we need data visualization.**

For instance, one of the things that I was explaining to the data science course is, what are the ways in which students attempt questions? We have keystroke-level data. So every time they type a Python program, every single keystroke, which second they solved which question and what they typed is available. And we had an agent analyze all of this, and it said some stuff like, "Look, there are four types of solvers, this is the performance..." Boss, I don't get it. Explain it to me in a way that I can explain it to someone who doesn't get it.

So it created this saying, "Look, in this almost 120-minute exam, students started solving this question, eventually got it right, then at this time went on to this question, then this time went on to this question," and you can see that it's an almost linear flow. These students are called the linear solvers. They go one after one after one after another. Sometimes they are not able to solve it and then they go on to the next one, and then they come back and maybe they solve it, maybe they don't, but this is their approach. In contrast with the cyclers. What these students do is they try this question, they didn't get it right. They come back to it. They try this question, they didn't get it, they come back to it, and again come back to it. They keep cycling back to the same question over and over again—that was a second kind of behavior.

Third kind of behavior are the jumpers. They skip questions. They don't even answer these in the middle. Fourth kind of students are togglers. They just get stuck with bad questions. So this student took almost 40 minutes on the first question, eventually got it right; took the second question, took another 40 minutes and eventually got it wrong; took the next question and came back to that question. But the point is they're moving slowly, getting stuck.

And if there was one advice, therefore, the TAs have to give the students... oh, yeah. Performance-wise, we could also correlate and say, **we know that the students who are going question after question after question—not necessarily solving them, but scanning them linearly—are performing the best.** Now, is this correlation or causation? We don't know, that's another evaluation. But we know that they're scoring better. So if the TAs had to give one piece of advice, it is: **skim every question at least before writing the code. Don't skip, don't toggle, get a sense of perspective.**

This is the most useful advice that we can give. But this advice would not have landed without them understanding the visual pattern by which students are solving. **In other words, if you want to convince somebody of your argument, you need to show a visual proof**, and the quality of that visual—the extent to which what you show matches what you want them to do—determines whether you succeed or not.

Another reason, therefore, which brings us to this, is not just explaining to other people, but actually convincing them of a point. One of the things I struggle to convince people about is how rapidly AI is advancing. So I use this as a visualization to represent this. X-axis is the cost per million tokens, roughly what will it cost a model to read all of the *Harry Potter*s or the King James Bible. And in June '23, Claude 1 could read it at about eight dollars.

The Y-axis is the intelligence: high school freshman, high school graduate, college junior level intelligence, and so on. And this is based on an Elo score; most people are convinced that we have some way of measuring model intelligence. And over time, the way the models evolved is that by November '23, we had a college junior level model, GPT-4. That was a big jump.

So think about it. In June, we had only a less-than-high-school graduate. But in a few months, we have a college junior. And a few months—almost a year—later, we have a masters student level. **In one year, four years of education is completed. And in four months, it's become a PhD.** And in a few months... okay, this GPT-5.6 [Taurus?] is just stuck there, ignore that. But in a few months, it has become a tenured professor.

So in about two, two-and-a-half years, from a high school student to a tenured professor. **That is how fast it is becoming smarter.** And when people see this, they say, "Ah, okay, AI is getting smarter at a pace that is faster than I had imagined."

The other thing they say is, "Ha, but AI applies in that field, not in my field." So I took... I had Claude take OpenAI's survey of a study called GPQA, which is: **across different professions in the US, where is AI doing better than experts?** Experts created a list of tasks, and humans were asked—human experts were asked—to solve it, agents were asked to solve it. Green means agents did better, red means humans did better.

Accountants and auditors: last year were doing better. This year, that is not the case. **My auditor submitted a return on my behalf; I had ChatGPT cross-check it.** I went back to her saying, "Why is HDFC, whose capital gains you've put in under Indian taxation, not considered under double taxation because I live in Singapore?" She said, "Oh, because it is an NRO account." I said, "The judgment does not say anything about NRO account." She said, "Let me consult my senior consultant." Next day she came back and said, **"Okay, you saved 14 lakhs."**

Fourteen lakhs! Are you crazy? And I know nothing about tax. I was simply asking ChatGPT, "Is she correct? Is she correct? Is she correct?" and gave it all my bank statements and so on. So at least for me, it beat—maybe not an expert, I don't know—but certainly beat my auditor. But for something like, let's say, personal financial advisors, even last year it was beating. Now when people start seeing this and look at the kinds of tasks that it's actually doing, they get convinced.

Now this next level drill-down becomes important for conviction. Because if I just said, "Oh, it is smarter," they'd say, "Yeah, but on what?" So I show them: this is an example of a task created by an expert for an expert. "Okay, 64%? Okay. And that was last year? Okay." Maybe not trying to move them from a 'no' to a 'yes'; if I move them from a 'no' to a 'maybe', that is a huge deal. But that is part of what we are using visualizations to do, which is convince them. And a list did not have that kind of an impact.

Finally, I use it to verify. Is what it's saying correct? One of the geospatial visualization analysis that I had it do, or one of my colleagues did, was a very interesting one. **I mentioned embeddings. Now, embeddings apply not just to text or to images, but also to geospatial data.** So I can take, for instance, the satellite imagery of one particular region as of some date and compare it with the same region, different region, same date, different date, whatever, and see how similar they are or how different they are.

So Varun had it look at various cities to see: are there specific regions that have become very different from each other? Let me see if I can get you a detailed version. So this is Gurgaon. And initially, the way it was constructed was he had a grid saying, "This particular latitude, this particular longitude, here is the difference in the..." whatever factor. Okay, this is very slow, but I know that it is also very... let it load. This may be worth seeing and I should have pre-loaded it if it at all loads. Any questions so far?

Okay, then let me just see if this loads up, it's fine. What we had, therefore, was a grid. But behind the grid, we put the satellite imagery, so at least I know what this place is. And let's take... which city? Do we have Chennai? Okay, yeah, we have Chennai. I don't know, I'm not sure how many of you are familiar enough with... okay, I'm just going to zoom in here and show you what Chennai's... okay.

So in 2015, this is what the satellite imagery of Adyar River looked like. This is what it looks like now. We say, "Hold on, wait, actually it's looking greener, but it's not like it didn't have that much water." But **if you look very closely, there actually is a lot less water in the Adyar River in 2015 in January than compared to 18 December 2025.** And the season is similar; between December and January, there isn't that much of a difference due to rain. And yes, this is a known thing that the Adyar River actually has more water this decade than it did last decade.

Another is somewhere in Anna Nagar, there was a massive concrete revamp. And what it was doing was taking the similarity... yeah, you can see that there are several new buildings that have been constructed that led to a loss of a decent chunk of greenery from the left. It was able to spot the difference beyond just the fading of the images.

There have been some places that have improved as well, like in Kodungaiyur. Where is Kodungaiyur? [Searches map]. So this ground has been converted to green land. Converted, or has become green land, I don't know, but there is definitely more green cover here. Positive news. And all of these were detected by effectively a grid overlay where it would say, "Look, I will take some particular parameter, like vegetation delta, and see where the vegetation delta is higher, where the vegetation delta is lower."

The reds indicate that there has been a significant reduction in vegetation delta. The greens indicate, if you see any green, where there is a significant improvement in vegetation delta. And I can see at a glance that overall in Chennai, things have generally worsened from a vegetation perspective. And each of those little boxes that you saw there as analysis was one particular cell where I could zoom in.

When we originally created this version, I had no basis for verification. **One of the first things to do, therefore, was to add a map behind so that I can see what this place is.** Then of course we gave it to *Times of India* and said, "You guys do the verification."

But they also said, "Can you provide us a closer look like this, showing the before picture, show me the after picture?" Have it do a Google search and an agentic search to find out what really happened in this particular area. And it starts showing some of these cities. So for instance, it's saying that this place in Bengaluru has moved away from its cramped legacy setup; there's a new BCCI Center of Excellence in North Bangalore and they have transformed 40 acres of rural scrubland into a high-tech sanctuary for Indian cricket. **Now you get context. Now I am able to believe it. This is a way of verification.**

Remember earlier for *Times of India*, I said, "Here is a way in which you say step-by-step, follow this process and you can verify, and anyone who processes it will know that it is true." That is one means of verification—that is how you verify the visualization. **This is the complement. This is where you're using the visualization itself as a means of verification.** You look at the picture from before, you look at the picture from after, you say, "Yeah, I agree, this has in fact changed. I can see the nature of the change and yes, this is valid." This is the other way in which, therefore, I use visualizations.

In short, here are the ways I use visualization: **one, to find questions.** What on earth don't I know in the first place? And this is something that you need to keep in mind when you are framing problems. There is the purpose and the framing of problems that we used to do three, four years ago, which is still important, but now you need to start asking much sharper questions on: what is it that I don't know? What is it that I don't even know I don't know? And what kinds of visualizations can expose that, keeping in mind that I can leverage agents to help me like crazy.

**A second is understanding and explaining.** Understanding for myself, explaining or convincing another person. And that is significantly to do with the design of the visual. **And by design, I don't mean the aesthetics, I mean the form, the function.** Ultimately, the person should take one look at it and say, "Ha, I get it. You are right." And use AI agents to critique your own visualizations—they're pretty good at it.

**Finally, to be able to verify at a glance**, whether it's for yourself or someone else, and that is functionally something that we will be looking at in great detail in the verification side. What I was doing in all of this was trying to give you a picture of three things. **One: agents can do most of the things that we have wanted to do, have been doing, in data visualization and can eliminate the need for decision-making and actions.**

**Data visualization in all of the cases that I was using it for, you'll find, was primarily for human consumption.** Where humans need to be in the room, either because we are the endpoint—we are the person consuming it—we are the actors, or we are accountable, we are responsible, we need to make a decision.

And if that is the case, then the purpose to which we apply data visualizations to needs to align with that. And you now have agents to help you in the variety of different ways that some of which we saw. And it is for you to discover those as we go along.

You will get stuck. The whole point is to get stuck. Here is one tip that I would ask that you try out: **anytime you get stuck, write it down.** "This is where I got stuck." I'll show you what my list looks like. I call it "AI Bottlenecks" and the most recent was a few days ago. "My scheduled tasks are not self-improving." I have a bunch of tasks that I run automatically; I want them to automatically improve and they're not. You'd say, "Boss, there are better things to get stuck on," but this is what I was stuck on that day.

"I have too many chats." In fact, I had to... I ended up making an entire application. The sole purpose of this application was to track my tabs. And as of now, I have 162 tabs open, which I manage across different windows and so on. And these are problems that I'm noting down because I'm stuck, and I'm hoping that AI will solve.

And some of these it's solving because, for instance... okay, that one's... yeah, too many chats. I had the same problem in July. But one of the things that I learned was that I should just focus on executing faster. And what that means is, whatever the tab says, give it to an agent and say, "You do what I'm supposed to do." Or close the tab, don't waste time on what's not clear. If I don't understand it, why do I have it open? So I ask an agent, "Do you think this is clear enough or it can be done in a better way?" If it says it can be done in a better way, I just say, "Log it somewhere," and close the tab.

And some things will get actioned anyway, some things will get duplicated, so I kind of found a learning. Some cases I've learned how to delegate, in some cases I've learned how to convert it to an asset. But the point is this: **at any point, I have a list of known problems, not vague unknown problems. Document yours.** And eventually come back to it, see if you can get yourself unstuck. But the more you get stuck, the more you have an asset because agents are problem solvers.

**If agents make the cost and time of solving problems down to zero, the difficulty is finding questions.** If you are stuck, fantastic, you have a problem to solve. Stop wasting that valuable resource. Every time you feel somewhat uncomfortable about anything or the other, please write it down—you have the best asset to pass to agents, and you need to train yourself towards this. Give it a shot, all the best. Any questions? Yes please.

**Student**: Sir, you talked about... you asked it and if it says it's not relevant, tell it to close or whatever. Does it apply to everything? So basically, I'm a person who won't be... I mean, I won't be satisfied if I don't go check it and then close it. I might think, "Okay, that might be something important," but the model says that it's not important. So how... [inaudible]?

My to-do list is 25 years old. Meaning I have items from 2001, and maybe slightly before, that I'm still storing. I won't throw it away. I completely relate to it. So don't—just archive it. Keep it somewhere. I've given up trying to train myself to discarding it. And the reason I'm comfortable saying "close tabs" is because I have it archived. I have another program which every day scans what are all the tabs that I have open, saves it, so I can always retrieve it.

**Student**: Sir, one more thing. So in the entire discussion, I think you didn't talk much about: **what if the AI makes mistakes?** How is... I mean, what if it says the correct tab is a wrong tab and tells you... [inaudible]. So in those cases, what should we do?

A very good point, and this is about how do we reduce hallucinations. I will answer this specific question and then generalize. So I don't have an agent closing my tab; I gave instructions to an application that an agent created that closes my tab. For anything that is moderately deterministic, I tell it: "Write a program, use that program and run it."

So I have, for instance, an "Edge Tabs" program which will automatically list all of my tabs. I have an "Edge Cookies"—this will list my LinkedIn cookies and so on. But the point is, **the agent, therefore, does not have to do something that may be non-deterministic when it can be delegated deterministically to a program.** So I tell it, "Solve the problem, see if you can do it through a program. If you're doing it through a program, save the program and reuse it next time." That's part of the answer.

Some things cannot be done deterministically and it may make a mistake. What I'm finding is that the number of mistakes that it makes reduces month after month with each new model for a given task. And then there are more complicated tasks that I ask and it makes a mistake on those. So what I do is benchmark it.

For instance, one of the things that I did yesterday—day before—was transcription. I have a whole bunch of transcripts; I said, "Try it out with GPT-3.5 Flash, try it out with..." [Searching] transcription summarization. "Luna," I think... Data stories... no, where did I... Luna, no, not here... Private search... yeah, summarize Luna, yes.

And basically asked it, "Look, which of these models should I go for: Gemini 1.5 Flash or GPT-4.0 Luna?" And the way it did it was it created a series of test cases. So it took a transcripts folder in which there were a whole bunch of transcripts, including, for instance, my daughter's conversation with a dermatologist. And when I said summarize it, Gemini 1.5 Flash summarized it in this particular way and GPT-4.0 Luna summarized it in this particular way.

Like this, I had it do it for every conversation. Then I had Claude Opus take a look at it—I would have tried Fable, but I said, "Let's not waste too much money on..."—and said, "Now which of these is better on what parameters?" So it ran a detailed check and found that between these, on the summary, Luna was actually better, keywords Gemini 1.5 Flash is better, blah, blah, blah. Luna is about half or 2.5 times less cost than Gemini 1.5 Flash, so it said, "Move to this." Next time I want to verify, I will do the same thing again.

**Is this perfect? No. Is this better than the other model? Yes, almost clearly and convincingly. Is this better than me doing it by myself? Absolutely.** I'm far more expensive and for this volume, I will make more mistakes. So **benchmarking is the other process.** You already saw an SOP for verification. You can tell another model to verify it. There are several ways in which you can verify.

Here is one prompt, and going forward when you get these questions, ask an agent. "Look, you hallucinate, how can I fix?" One of the prompts that I had given it—and I know we are almost out of time—is, "**How am I going to verify you when you are smarter than me in taxation?**"

And it said, "Anand, this is a problem that has been there for centuries. How are you going to verify your auditor who's smarter than you at taxation? How do you do it? How do we do it? How do judges pass judgments on a patent when they don't understand engineering? How do regulators who don't understand telecom pass a telecom bill? **We've been solving this problem for centuries. How does the FDA verify drugs? They have a checklist of a process.** If you've followed this process, they deem the drug to have been at least not too much of a disaster. So following a checklist is one of those standard procedures."

"Second, how do we make sure that the real estate agent is not going to cheat me? Because he's going to get paid only after I get the house. Outcome-based pricing or incentivization, commission—that's another." There are so many techniques that we can use. And the prompt that I gave it, therefore, was: "How can I learn from the wisdom of all these other professions in the past where people have verified people smarter than themselves, and I'm applying those techniques?"

You probably will have a bunch of other questions. Feel free to drop me an email. Just search for my name, S. Anand—Google me, you'll probably find me as the first or second hit—and yeah, my contact details are there. Just welcome to drop an email anytime. The second link is me. The class is technically over, however you are free to hang out. If someone wants to leave so that you have other classes or something, please feel free to do so. But if you have other questions...
