# Data Stories with AI

**Anand**: [00:01] Hello, good afternoon. We will be starting in just a bit.

**Rohit**: [00:05] Hi Anand, good afternoon. Must be what, 6:00 or something?

**Anand**: [00:14] 3:00 PM. I'm in India.

**Rohit**: [00:15] Oh, you are in India. Okay. So, you are in our time zone, which with you, it's difficult to know where you are. So, our time is 3:00, right? So, maybe we should give five, six minutes because people will be just getting in. No, I was thinking, Anand, about this Mahabharata—I mean, that was one of your first few works I had seen when you had visualized. I still remember, and if I were to open it, I'll be blown over again. I had no idea that an 18-book mythology could be, if I can use the term, "data-fied" and visualized the way you did. And this was 15 years ago. So, those works still exist, right?

**Anand**: [01:04] The books, definitely.

**Rohit**: [01:05] No, you had done that character linking of Mahabharata. I remember very well. That was one of your first works, maybe 15 years ago. I am forgetting when I had seen it. I couldn't believe what I was seeing then.

**Anand**: [01:23] I’ve just put the link in the chat. The same one; it's still there.

**Rohit**: [01:27] Oh yeah. Oh, I still remember. I of course have many of your links. Oh, yeah, this is exactly what it was. I just couldn't believe it the first time when I had seen it. And I remember spending hours because, I mean, I don't think there's a book—or I'm sure there are, but I don't know of—with so many characters, so many events interwoven. And I sort of—and this is of course, I had not known of AI then. So that, and I think Parliament stairs. It’s just the power. And people who are on the link may want to just briefly, as we wait for others to join, click on this link that Anand has shared, which is the work of Anand I had seen in 2011, I think, or maybe 2012 max. So, which is 14 years ago. And so this was, and obviously this was a done job. So obviously you would have done it much before I saw it. So, the scale of data, the thought, the visualization, I thought was quite a thing. Sajeev, you had seen this? No, I don't think I ever mentioned this to you.

**Sajeev**: [02:50] No, I haven't.

**Rohit**: [02:51] Yeah, then this was—I was telling Anand, in my recollection, I was trying to think of the kind of—hi Tanay, I finally see you.

**Tanay**: [03:08] Hi, hi Rohit. Hi Sajeev. Hi, hi Anand.

**Rohit**: [03:13] And this was—and I also remember, Anand, you had developed some, your personal movie selector or something? You used to—there was some tool you had made to help you and your family decide which movie to watch. And I thought, you know, that was such a, as they call it in the startup world, a problem to solve. These are all, Sajeev, Tanay, whoever is joining in, this is all about 14, 15 years ago, Anand's work that is coming to my mind. And I remember, you know, I didn't know "C" of coding, as I still don't know coding, but the possibility of it had really... and then of course we did many more works together. Sajeev, how many had registered? We will wait for, I think, another...

**Sajeev**: [04:14] We have around 250 people registered. Oh sorry, 300 now.

**Rohit**: [04:19] Oh, wow. Okay. That's quite a—it’s wonderful to see so many people interested.

**Tanay**: [04:28] Yeah, Rohit. It’s—I mean, I think it's the field that we can't have enough of. Even if people who don't ever want to do data journalism, they just get excited about data, the whole of journalism will benefit.

**Rohit**: [04:52] You must be... Sajeev, Tanay is coming just a floor above. When are you coming, Tanay?

**Tanay**: [04:57] Monday.

**Rohit**: [04:58] Oh, that's just day after. Okay.

**Tanay**: [05:00] Yes.

**Rohit**: [05:02] Good, very good. Anand, Tanay was—is joining Economic Times. Tanay, I'm taking the liberty to tell.

**Tanay**: [05:09] Yes, yes.

**Rohit**: [05:10] And he was heading Mint's data team. He's done many other things before that. But this was his... yeah, I know you from the Mint work, Tanay. Congratulations.

**Tanay**: [05:22] Thank you, thank you.

**Rohit**: [05:40] How long, Anand, did Gramener—Gramener, which obviously—so, when was it founded?

**Anand**: [05:48] 2011.

**Rohit**: [05:49] 2011. Okay. I mean, many of you may not—some of you, many of you know of it. Gramener's name was the initials, the first names of the founders. Is that what the case was, Anand?

**Anand**: [06:05] Also. Also, we started as a rural BPO and energy management company, and "Gram" Ener. But it also happened to be an acronym for all the founders' names.

**Rohit**: [06:17] Okay, okay. Gramener was one of the first, at least in my knowledge, of this scale of data visualization and pattern recognition that I had seen for the masses. You know, we’ve known as some of us, those who are journalists here, and I think many—you know, we've heard of Mu Sigma and the huge unicorn data, but that data had nothing to do with—it was a very B2B. Our interest was more as a corporate entity than any potential user. Gramener was my first exposure where, you know, they were doing this cool stuff which we wanted to—you know, from number of hot days in a year, to Mahabharata, to movie selector. Of course, many, many more serious things. I think one best cities survey of India Today also Anand’s team had helped. You know, **city ranking is a very difficult, complex job because databases are very poor.** States—those who are here, if they consider state databases to be poor for ranking, which it is, you try ranking cities, and then, you know, all hell breaks loose. You have no idea where you are headed because...

**Tanay**: [07:31] And just the different kind of formats that each city’s data is in.

**Rohit**: [07:35] Yes, yes. And see, people, your readers, relate to cities more than states. States you can get away by saying many things. Well, you shouldn't, but I'm saying hypothetically. **Cities, people react very strongly because they relate with the cities more than they do states.**

**Tanay**: [07:51] Janagraha has done great work recently on city—for municipal data. So, they have a portal called cityfinance.in. So, they're trying to standardize all, you know, all municipal data across all cities in one common format and also making it queryable now.

**Rohit**: [08:12] Oh, yeah. Yeah. But you know, Tanay, the Anand thing, it had—it had so many things. At that time, from number of cinema halls, to commuting time, to commuting cost, to ease of finding a house for rent. You know, all—I mean, we were very ambitious, but Anand did not curb any of our ambition and he said we will get everything in. I remember it was quite an exercise, trying to rank cities, and then we did the existing... so, in many of these things are still not done well, Tanay. I mean, there's no reason why India should not have a live city ranking, you know, for the common person, a database and visualized dashboard. It’s so useful, right? For... so, you know, these are the things I think we should really think seriously. Anand, we can start whenever.

**Anand**: [09:12] Yeah, I think we should start. Rohit, you can...

**Rohit**: [09:14] No, thanks. I'll just—thanks everybody for joining. Many of you may not know—Anand told me his GMeet can take 1,000 people participation. So, I was fine, because Sajeev was telling me the number of registrations, I was getting a little worried. But thanks so much for your interest in this. Many of you don't know me, but that's perfectly fine because I'm not going to do anything other than just introducing Anand.

[09:47] Some of you who have been here since 3:00, you know my friendship is how I describe Anand. More importantly, **admiration for his work is at least 15 years old.** As I said, for the work that comes from—and I've seen so much of his work that he probably doesn't even know. But when he had visualized and what I would call "data-fied" Mahabharata, and I was very impressed when he shared the link. He used to make cool stuff like how to select a movie to watch with your family every weekend or whatever. And then he helped India Today Group and, I think at that time, Network18 or CNBC, do election visualizations for website and television very early on. So, **he has been doing large-scale data work for I don't know as long as many of us have not even known and done.**

[10:37] Anand, one other work I had seen of you of CBSE names and date of births and patterns on date of births, why certain months certain kids are born and names are given. So anyway, he has been—I’ve been seeing this work much before I knew anything called Natural Language Processing or GenAI. So, Anand has really been a trailblazer in this regard and he's been very generous with his time and willingness to share all his knowledge with everybody.

[11:15] So, he's—I'm so happy. And I'll just end by saying this thing about AI agents, and Anand will explain better. A few months ago, when I met Anand at a hotel reception and among the first thing he told me, which totally threw me off, he said, "Rohit, let me hire four, five people for you. Tell me where is your skill gap." I was taken aback. I thought there's some deal, secret deal between my company and Anand to hire employees for Times of India. But then I realized **he was actually talking about AI agents and what they can do.**

[11:46] And that's how—and then right now in Times of India Print, there's a property going across editions, which I think is a—one of the rare examples of being AI-augmented and AI-automated both, and I would say even AI-customized. There's a property in TOI Print that Anand has almost solely driven and partly in the last mile helped by Sajeev, my colleague here who is in this group, done. And what it showed us—and I think that's very important for all of you to know—is that **AI agents have reached a stage where anybody can benefit to do any kind of work you want to do.** It doesn't need to be... and at least that's what Anand's work showed it to me.

[12:30] **It will augment parts of work that you want to augment and you needed or thought you could augment but you didn't have time or patience or skill.** It will automate parts of your work that you less like or didn't want to do but had to do. Most importantly, I also learned that it can customize it for you. So, it's not that AI will do something that is out there in, you know, and anybody can do it. So, **you can actually develop unique competitive advantage, augment your skills, and automate tasks that you don't want to do.** And this is just in a short period of two, three months. I'll just stop here and Anand, thank you so much and you can take it over and set the rules of anything you want to.

**Anand**: [13:16] Absolutely. Let's dive in. And I'd love to start with what you had just explained, Rohit, which is how we went about part of the way Stat-o-istics and—I’ll also show—geospatial insights and so on. But this can be a workshop, meaning you can—each of the agents are in your pockets as much as my laptop. And therefore, let's start with something simple. What I'll do is share a link in the chat that we can all fill out, and I'll also share my screen so that we can see the contents. Okay.

[14:06] So, this is the link which I will also place on the chat window. Or yeah, and all you need to do is visit forms.s-anand.net and start—and this is a live form. I’ll be adding prompts here, I'll be adding questions here, we can all work on this together. Because what we have is an opportunity for several dozen of us to create data stories on the fly using agents, possibly publish them, who knows what we find. Let's do that.

[14:43] A few simple questions. I'm just curious what you do. I am a researching—hello, I'm Anand as part of Straive. I co-founded Gramener, a data science company before this. Rohit gave you much better introduction than I just did. I've been creating data stories for many decades, some of you may have been doing it for more, less. I'm right now in Chennai and these are independent answers. You can just submit an answer at any point, you can't edit it. And today, I am going to be using primarily, I guess, ChatGPT. I know that I'm giving only one option. In reality, people will be using multiple options, but for today I'm primarily going to be using ChatGPT, though it could have just as easily been Claude.

[15:43] What I would invite is for you to pick any one tool—of course you're very welcome to use multiple tools—but before we get in, if you said, you know, "I'm going to pick this tool to play around with data to tell stories," what might it be? **It helps a lot if you have a paid subscription because then you can get it to run longer, but more importantly, use smarter models.** And the difference the smarter model makes and running longer makes is a lot, and you will see it.

[16:13] How confident am I about getting AI to find a real insight in data? I have absolutely no doubt that today it's going to blow my mind because it does that every time. I've done this so many times. Now, the problem is not that it blows my mind, but the problem is that I'm not able to absorb the stuff. So, it's a very different take. Maybe the real question that we need to answer is, **if we have so many insights and they become so cheap, then what do we do? How do we distribute? How do we prioritize? Do we act on all of them at a time?** The problems start becoming very different, and I don't even know how to frame what the new problems are. But as I said, good problem to have.

[17:01] I will be adding more questions as we go along, like what problems might you want to work on etc. But for now, let's start with what Rohit had shared, which is something that my agents have been working on. The starting point was the "Hack of the Day" series, which is a series of crisp life hacks that you can apply that are technology-oriented. This has now led to a site—and I'm going to actually just start putting in the sites as I visit them on the chat window, so you are very welcome to browse them. Now, the page that you're seeing on the screen right now is AI-generated. And what I mean by AI-generated is in two parts: the content and the representation.

[18:07] I'm going to focus initially on the content. Generating an interesting life hack is easy. Once you've done that a hundred times, then it starts getting boring for the people who know enough about life hacks. It gets handed down. And then people start struggling to get enough good quality—you know how series evolve. And therefore, what I asked my agent to do on the first day was—and when I say "my agent," I'm just talking about ChatGPT.

[18:49] So, Rohit had shared a bunch of images saying, "These are what we publish as hacks of the day," little cards like this. **One of the first premises is, if you have an agent reasonably smart working for you, you should get the agent to do the work.** Normally, what I would have done, if I wasn't practiced in working with agents, was I would tell it, "Look, a hack of the day will typically contain a title, it should contain technology-related stuff, it will have a 'what it solves,' 'what to do,' blah, blah, blah." Why should I do that? It should do it. What is an agent for?

[19:31] So, I said, "Analyze these. And if I had to ask an intern or an AI agent to create a bunch of these, what prompt should I give?" Now you'll notice that I'm not asking it to give me another hack like this, or even a bunch like this. I'm doing what is now commonly called **"Meta Prompting."** Why? Pure laziness. I have to think of a prompt. I'm giving you examples. I'm going to ask you to do something. You tell me what I should be asking you to do so that I can ask you to do it, and then you do it, and then I will ask you to check it. **Basically, delegate all of the work.**

[20:12] In the last few months, this single philosophy has probably been the most effective for me, which is to **watch closely where I feel stuck and just write it down, saying, "Here's where I'm stuck today," and copy-paste that into the AI.** At the very least, it might get me unstuck. But then one of two things happen: either I get unstuck, or I get stuck because I have to sit and read its response, which I may not understand or it's going to tell me to do hard things. In which case, I write that down again. "Now I'm stuck. The response is too complicated, too difficult to follow." And then I copy-paste it back. Until I either get unstuck or bored. But **the results have been one or the other. I have not stayed stuck. The bottleneck has constantly shifted.** And that may be the defining characteristic of what AI is for.

[21:18] But now we've gotten into philosophy, let's get into practical stuff. So it said, "blah, blah, blah," and I'm not going to go through what it said because I didn't go through what it said. Ultimately, I wanted a prompt. It gave me a prompt. So, I took this prompt and gave it back more or less, and said, "Okay, here's a better structure." Fine. I just went all the way to the end. And ah, I like this. "If I want to mimic the Times of India style more closely..." But the reason I like it is because it feels smaller. Anything that's small is easy. Some more stuff I mostly didn't read all of this.

[21:54] Next, I said, "Do some research. Maybe I gave you too little information. Maybe there's additional information. You go do the research and see what else is out there." And it went through recent hacks, it listed some additional hacks, it found some of the hacks that were there, and it made a set of observations. Looks like community digital services are there, banking and financial services are there, blah, blah, blah. And it can combine a—and it says, "I can create a list for you." Great.

[22:25] Now I said, "You give me more." You'll notice that I did a copy-paste of the prompt it gave me. In a sense, this is sort of like telling an assistant, "You go read all these documents and summarize it for me." The summary is really not for you; the summary is for the assistant to have gone through it and understood it. Now I know it's there. I can copy-paste it, but it can copy-paste—or he can, or she can copy-paste—as well as I can. What's the point? Now you just give me ten more.

[22:52] And it identified a bunch of these, which I didn't read. I said, "Give it to me in the format of these cards." Now I got this JSON by copy-pasting from somewhere else. Now, why in the form of a JSON? It's a little more structured. What is JSON? Some structured format. You could have asked for Excel; doesn't matter. Ultimately, we want some structure so that I can sort, I can search, I can filter, that sort of a thing. And it gave me these cards.

[23:22] I passed it to Sajeev, Rohit, and others, and they basically said, "Look, content looks fine." And then we went on into the format. There was, I think, only one tweak, which is that it needs to be more technology-oriented, not general hacks. So, other than that, this came out. And since then, I've been having a series of conversations roughly every few weeks. Here's another... and you can see from the nature of conversations: "Nice. Continue research extensively. Yeah, continue them in the same format. Yeah, this is nice, but some of these have already been published. Yeah, continue the same. Here's some feedback I got. Continue, continue, continue." And if I wanted to generate another dozen more hacks now, let's do that.

[24:14] So, here is the most recent one that I executed. I'm going to just copy-paste this again. It says, "Identify 20 more interesting ones to include. Don't repeat what you've already mentioned earlier. Use the same format." And last time I added, "Look, think about it. What are some of the kinds of stories that we've uncovered? What haven't we covered? What will be of interest to this audience? Where should we check for such sources? How can we carefully evaluate what will be most effective?" In other words, I'm giving you some food for thought. Plan and then execute, and then make sure you fact-check everything and share in the same format.

[24:56] Let's run this. It might say this conversation has gone on for too long, or it might continue, who knows. We'll figure it out. But this is it. **This is what it takes to automate an entire property: asking ChatGPT to do this.** And you'll notice that the continuation of this, once it's been set up, was simply copying something and pasting something. So, effort is really one minute of my attention and about 20 minutes of its effort, which I don't care about because I can always just go on to another tab. This was simply "research, find interesting stuff, and give it to me in a certain format." And that in itself is powerful. But is it data? No, maybe yes, not sure. Borderline. Let's assume maybe it is not quite data. It's research stories that we have.

[26:02] Let's do a quick poll. Can AI actually process data? So, I'm going to add a question here. Let's call it "Hallucinations." And the question that I'm going to add is, "**How confident are you that AI will give correct answers to numerical questions that data involves?**" And yeah, let's give it a set of choices. The same link that I had shared with you—which I’ll put on the chat again in case someone's joined afresh—should now have, yeah, as question number six, the very last question: "How confident are you that AI will give correct answers to the numerical questions that any kind of data analysis involves?" Give it a shot.

[27:01] If you say 0%, you're actually very confident that it will not give correct answers, it'll hallucinate on numbers. And what I mean is real full-fledged data stories. 100% means that you're sure that it will not hallucinate, it will give the right numbers. I'm not filling in my answers just yet, but I'll just do a quick check to see how many people have filled this out. Okay, we have 19 responses so far. I’ll give it until we get about 25 responses. Okay.

[28:03] To recap, I'm just suggesting that everyone take a shot at answering this question on the link that is in the chat: "How confident are you that AI will give correct answers to numerical questions that data involves?" Now, the reason this is important is, **if AI is going to mess up on arithmetic, then at least that part of the work you should not delegate to AI.** And a big part of knowing how to use AI is about knowing where it does a good job, where it does a bad job. So we take the good and drop the bad.

[28:46] Okay, yeah, thanks, we'll just keep the responses continuing. What we'll do is, let me do one... I'll tell you a quick check that I ran on how well LLMs are able to do multiplication. I asked 50 models to multiply two numbers: two-digit numbers, three-digit numbers, all the way up to nine-digit numbers. This was about a year ago. And here are some of the results. Let's take a model that would be a good model: Gemini 1.5 Pro, which was released in March 2024. Now, this couldn't even do three-digit multiplications properly. One out of five times it got it right, four out of five times it got three-digit multiplication wrong. Luckily, it also I think got five out of six digit multiplication correct.

[29:46] Some of the other models were better. Among the better models were Claude 3.5 Sonnet and Opus, which managed to go as far as five-digit multiplications. Any of these managed to go as far as five-digit multiplications. But then some of the...

---

**Anand**: [00:00] ...and then saw the newer thinking models come. They were able to go as far as seven-digit multiplications, but not nine-digit multiplications. Now, we have not one nine-digit multiplication, but in data analysis, we'll probably do a thousand. So, there's a good chance that it will go wrong.

**Anand**: [00:21] However, people figured out something clever—that **you don't need the model, which is a language model, to do the multiplication.** That's like asking a person, "Can you multiply two nine-digit numbers in your head?" When a person doesn't think in mathematics; computers think in mathematics, people think in language, roughly. Instead, what we would do if I were hiring an accountant was give them a calculator or a computer and tell them to do the job. That person will say, "equals this cell into this cell," effectively doing an invocation, calling a tool that will get the job done.

**Anand**: [00:59] And all the AI agents started following this principle. And the magic words to invoke this are simply to say, "Write a program to do X. Write and run a program to do X," which on ChatGPT can be as powerful as we like. And it's not even just simple arithmetic. Actually, let's ask... let's ask it to do something.

**Anand**: [01:34] "I'm in a workshop, and we are exploring the capabilities of large language models to do arithmetic, keeping in mind that large language models don't really do mental arithmetic well, but they can write programs that do powerful arithmetic. Can you write and run a simple yet mind-blowing program that will help the audience understand the sophistication to which LLMs can reliably go? Keep in mind that my aim is that the audience will appreciate the following points: First, it is simple enough for them to glance at it and understand, and this is a lay audience. Keep this to the level of something that, let's say, a grade eight student can understand. Second, it should be obvious that this is something that is both useful as well as difficult, one would think, for a language model to be able to do based on a conventional understanding. And the instant I look at it, it should come through this way. Think carefully about the objective, but make sure that what you finally present on screen is concise."

**Anand**: [02:57] Now you'll notice a couple of things. First, that I was talking to it. I find that very convenient, especially in workshops, because I'm talking to you as much as talking to it. Second, that I'm rambling a bit. If I knew exactly what I wanted, then I would have given a far crisper brief. Maybe I would have told it, "Write a program to do some specific thing." But **talking helps me think. And if I have an assistant who is infinitely patient, I may as well use that patience, use that conversation to sharpen my thinking as well.** In other words, when you're crafting data stories with AI, you don't necessarily need to know what data stories you want to craft. The story can evolve along the conversation, especially because the capabilities are increasing steadily. **What you have in your pocket is not just an expert language model; it is an expert journalist, it is an expert programmer, it is an expert financial analyst, it is an expert bureaucrat, it is an expert anything that the world's knowledge can possibly encompass.** There are many things that it is not. It is not an expert parent. It is not an expert... I was going to say empath, but no, it actually has a fair bit of empathy, so I take that back. But a bunch of things that it is not. And what it is not is really the bulk of what I'm trying to figure out.

**Anand**: [04:35] But okay, it's saying: "8 foods created 255 possible baskets, with 20 foods..." what? Okay. It's saying an LLM may not reliably multiply big numbers in its head, but it can write a tiny program that checks every possibility exactly. That's a good point. This is arithmetic as a tool, not arithmetic as memory. And what it seems to have done is... let's see what it did. Okay.

**Anand**: [05:21] Now, here is the problem: it thinks it's done something, but it has done a pretty poor job of explaining what it did. So, "Explain to me in very simple terms what you did. Just two or three sentences covering why this is interesting." And this is another principle that I will share, which is: **do not give your assistant the respect, the time, the space and all for them to explain in great detail what they did. Cut them off. Tell it to me in terms I can understand. I have only ten seconds.**

**Anand**: [05:46] So, it gave the computer a menu, a budget, and a calorie limit, and made it try every possible food basket and picked the one with the most protein. Aha! And to be fair, maybe it explained it somewhere above; I didn't have the patience. **And this matters because it doesn't need to do the arithmetic in its head; it can write a program that can do the arithmetic and check hundreds or millions of options reliably.** This may be the single biggest enabler for our session: the fact that large language models can, in fact, write programs reasonably reliably and get to something that is insightful. Which will lead me to my answer for this question. I'm going to put this at 90%. Why not at 100%? Sometimes I'm not sure if I'm asking the question right. Sometimes I'm not sure if, even though it doesn't make a calculation mistake, it's understood my question correctly. But 90%, maybe more, but that's where I am personally.

**Anand**: [07:00] So, let's now dive into the actual "doing stuff" part where what we'll do is go through the stages of finding some data to analyze. And let me ask you, what is a data set or story that you would like to explore or write about? And let's pick something concrete. This question should be appearing as question number seven in the same form, for whatever it's worth. Put it on the link again, in case anyone's logged out and logged in. But just go all the way to the end and find question seven.

**Rohit**: [07:31] Sorry, Anand, one quick question. Should participants now, at this stage, write here what, however vaguely or clearly, what is it that they want, which data, which field, which area they would like to do in today's workshop? Is this the part?

**Anand**: [07:49] Yes, please. I would like all of you to fill in this particular one because **what I'm going to do is take all of your responses, feed them to ChatGPT and say, "This is what my audience wants. Find me something that satisfies the majority of them."**

**Rohit**: [08:05] So guys, please, whatever may be your area. It could be food, it could be aviation, it could be, I don't know, India's fertility rate, it could be anything.

**Anand**: [08:17] And once I have maybe 20 responses, that should be... 25, let's say. We got 35 on the earlier one. So, once we have up to 25 responses, I'll collate all of those and put them into ChatGPT and use it for the first phase, which is data discovery. Meaning, I don't want to sit and discover data; you discover the data. Incidentally, I'm in a sense using you as agents—the audience. I'm delegating the task of, or crowdsourcing the task of what is interesting for this audience.

**Anand**: [08:52] But I can also ask ChatGPT to do this: literally taking the... all of you would have probably filled in a form. Taking that Google Form, I would say, "Research each one of these people, and then act as if you are that particular person, and then fill out the answer to this question: what might they have said?" It turns out that it tends to be 80% accurate. Well, it may not know your persona well enough just based on a Google search, so we may have to fix our query. But for the ones you know... there are many famous people here. So, it will find out, it will impersonate them mentally and answer the question. And I can... that's another form of crowdsourcing as far as I'm concerned.

**Anand**: [09:43] We have almost two dozen answers, getting close. So, let's see. Almost there. One more, and I will take what we have. Just waiting for one person to click the submit button. Take your time, no pressure.

**Rohit**: [10:05] Yeah, people must be thinking what to ask, and which is fair.

**Anand**: [10:09] We just crossed that top, so let's take all of this and... I have a bunch of... you're probably wondering what on earth I'm typing. Ignore me. It's the least important part of this stuff. So, now I'm going to share.

**Anand**: [10:32] "Look, I want you to find some interesting data sets. Specifically, I asked the audience of the workshop that I'm in what kinds of data sets or stories they might be interested in having AI guide, and here are the responses that they've shared. I'd like you to find data sets that satisfy as many people as possible. Don't give me just one answer; give me multiple answers so that I can pick from them."

**Anand**: [11:04] This is the first part of the prompt. Second part of the prompt is simply your answers pasted, which... okay, yeah, all of these. So, I'll just call this "answers," and you notice that I'm putting this in brackets, in angular brackets, just so it knows that that's the answers string. There's one more thing that I'm going to do, which has come from some amount of experience, some amount of experimentation, which is called the **"Ideation Protocol."** I'm going to tell it specifically how I want it to think about this. Let me explain what that prompt was about.

**Anand**: [11:43] Actually, I'm going to show you two things. I'm going to copy this not just to ChatGPT but also to Claude and have opus—no, in the default thinking mode, run the same thing. Now, this prompt is a brainstorming prompt. **When I want it to brainstorm, I ask it to diverge then converge. Diverge by coming up with a whole bunch of interesting ideas. Find people who will see this differently. Find domains who will see this differently. Then pick the obvious ideas and get rid of the obvious ideas. I don't want standard stuff. And then merge all of these and then generate lots of possible ideas and then pick from those which ones are likely to work, which ones are likely to be interesting, etc., and then give me the result.** This will take a long time.

**Anand**: [12:44] Now, there is another... so what this was geared towards was finding data sets that are interesting. I have another lens that I want to gear for for this workshop. If I had three hours, two days, whatever, I'd say, "Okay, let's go for the best." But I'd like to do this part of the exercise in another, max, half an hour. So, the change that I'm going to make here is not this but, **"Make sure you find a data set that I can download as a single file under 10MB, is popular, public, easy to process, etc., in a workshop."** This is important. Okay, I will not say "this is important" because ChatGPT tends to already follow my instructions too well. If it were Claude, I'd have told it, "This is important." With ChatGPT, I'm too scared because if I make even the slightest mistake in my invocations, it tends to follow that even though I might be wrong. Claude will tell me, "No, no, no, look, I know what you mean, I'll get your job done." ChatGPT will say, "No, I know you said this. I will go to the end of the earth and solve this problem for you." Use them to their strengths.

**Anand**: [14:09] And let's spin this. Make this a bit bigger. At least one of these will come in handy. There may have been more data sets that you have... yeah, you have five more, six... four more responses, but that's okay, this should be reasonably representative enough. Now, you notice that the data set discovery problem is, at least in one sense, easy. **There are two kinds of data set discovery approaches: One, I have a data set in mind, I want to write stories out of that. The other is, I want to write stories broadly in some area, maybe not even in some area, what can I write about?** And that second problem—that of discovery—this is good for. What we saw in Hack of the Day was that kind of a problem. "Think broadly. I want to solve some tech... and I want to come up with a tech hack. So, research and find out what would be appropriate." This is similarly that "research and find out what is appropriate" kind of a situation.

**Anand**: [15:21] Note that there is no cost or danger in running this. If it gives something good, it's a bonus. If it gives us something that either doesn't work or wrong or doesn't get us something, we lose nothing. In other words, **there are several low-cost situations where we can apply AI. And that may be the easy boost. Why bother taking a risk? Start with what's easy, find out what works, then scale.**

**Anand**: [15:47] So, now it's thought through as a data journalist, as a sports wonk, as a service designer, investor, etc. And it's also doing some hospital triage and fermentation, whatever that is. This is applying all the fancy thinking that I told it to do. And voila, it's... okay, let's take a look. One data set is "Oil to Wheel to Wallet." It's saying take Vahan data by fuel category or state, SIAM, FADA sales, PPAC, NSE sector data, etc. And what is it going to do with that? Somewhere it would have explained, or will be explaining, what the whole "Oil to Wheel to Wallet" story is about. "India starts shrinking before India shrinks." Oh, fertility, population shrinkage, life expectancy, which can come from NFHS, SRS, etc. Fair enough, it's coming up with story ideas, which is a good thing. And it's saying that this is the best bet.

**Anand**: [16:58] Something that I would love to explore for tomorrow, not for a half an hour quick burst of "let's do something with it." Which is why I'm going to take the second one, which is still running and... okay, it still seems to like the whole vehicle registration data thing. Which leads me to think maybe the Vahan data is not as difficult to download. And okay, there's a data.gov.in link. But doesn't that require a subscription? "If anybody from data.gov.in is here, we can happily give you a subscription." Let's see if we can get the part... and seven with the bit...

**Rohit**: [17:42] Anand, should people ask the prompt that they have put you here on their own laptop too, or that's not advisable?

**Anand**: [17:49] The prompt that I shared was a collation for all of those. The prompt that you shared is simply "Find me a data set for..." copy-paste what you typed as a response. And we, in fact, should do that. So, yeah, no, please, this is an invitation. Please give it a shot now just to see what you will find. I've been bugging you for 45 minutes; we will soon switch over to you doing a fair bit on your machine. In case whatever data you have shared with Anand for him to find out... I suspect maybe some of you would want to put on your own AI and see what you get. So Anand can probably then tell you, right Anand, where...?

**Anand**: [18:35] Absolutely. On your GPT, Gemini, Claude, whatever you may be having. Exactly.

**Anand**: [18:49] So now, my next question is... to be fair, this is a 299 rupee download. I can just put in a UPI and do that. But what if it is able to find me something where I don't have to do it? See, the good part is, while the data.gov.in data might end up being far more cleaned, proper, etc., **I want to use this opportunity to show how we might be able to take unclean data and have it do the cleaning.**

**Anand**: [19:19] So, okay, that is this, and it's saying "Source from blah, blah, blah, Last single etc." And... voila. So it is saying that is currently the best source. Why not? Let's just go ahead and I'm going to download this data set. Let's give it a shot. And scan... SVG code... okay, payment's made, it's processed, payment successful. And let's download the zip file that contains... let's see what it contains. Actually, I'm not going to see what it contains because I'm going to ask my assistant to see what it contains. In fact, all of you should ask your assistants to see what it contains.

**Anand**: [20:47] Let me open this. What I'm going to do is put this somewhere where you can take a quick look. Create folder. Go back... a little... best place. Put it in a minute. I will... it's not very large. How large is it? Not very large. Okay, 11 kilobytes. Okay, that's suspiciously small. Hold on. Okay, no, it is actually just 13 kilobytes. No, that is too small. That is not worth analyzing. 13 kilobytes I can do in Excel.

**Anand**: [21:37] Let's find something bigger. Let's see what else it suggested. Okay, "World Cuisine Ingredients CSV." There's something that's around 214 kilobytes on Kaggle; that's reasonable. Let's take World Cuisine, and the fact that it suggested it means that somebody here is interested in cuisine. Let me download this. And even this is small; 135 kilobytes. No, it's slightly larger. Okay, 69 kilobytes, 49,000... "International Football Results FIFA," slightly more promising. This looks reasonable. "Open University Student Assessments Data." "Delhi Temperature." "India Crude Oil Exports."

**Anand**: [22:31] Okay, let's just go with the... Rohit, I'd like you to take a pick based on what you see on the screen. Student Assessment Data, or Football Results, or World Cuisine?

**Rohit**: [22:47] Cuisine... yeah, I'm guessing people here have asked for these, right? But and you're looking for large. Do you want to look at... some are interested in AQI, but I don't know how large this database is.

**Anand**: [23:05] Let's find out. Actually, sorry, I really shouldn't be doing this. I should be asking it for something reasonably large. And okay, yeah, this looks good. Now, there are a bunch of CSV files. This is not very clean. Question to anyone in the audience: **supposing you have something like this and you have to download a whole bunch of CSV files, how would you do that?** You can unmute and answer, you can type the answer. Anybody. A bunch of CSV files on a page, clicking on it downloads it. I just want to download the whole lot. I'll give it ten seconds to see if anyone has an opinion.

**Tanay**: [24:06] Anand, can you hear me?

**Anand**: [24:07] No, no, you are... you are banned.

**Tanay**: [24:14] No. Claude has a browser connector, so I'd pass on the URL and say "download the CSV files from this page."

**Anand**: [24:22] Claude on the browser? Yes. So there is a Claude extension. Now, it probably won't work on my machine because I'm on a Mac as part... let's give it a shot. It's got a bunch of models. So, this is the Claude extension for the browser. And you can choose the models there. But the interesting thing is that it can see what's on my page. So, "Zip and download all of the CSVs on this page." And... let's see. It's going to act without asking. Fine. Yeah, let's just run it with Sonnet, which is a medium quality model, and see what it does.

**Anand**: [25:10] So first, it took a screenshot. You may not be able to read this very clearly; I'll just walk you through what's going on on the right side. It took a screenshot, it's extracted the page text, then it's saying something about the JavaScript tool, something else about the JavaScript tool, and... it seems to be acting on the page. But you notice one little red arrow out here. That is basically the Claude plugin taking control of my browser page. If needed, it can go click on each one of these. But hopefully, it will be smarter than that and do something more efficient than having to download one by one and then consolidate. It'll probably be able to download the whole thing at one shot.

**Rohit**: [25:57] Anand, for your benefit, so that people here know, is this Claude extension a paid thing or it's a free...?

**Anand**: [26:05] **I think you can use it even with the free version, but you are likely to run out after maybe one use a day kind of a thing. And therefore, you should anyway pay for one month, $20.** Try out all kinds of stuff. And if any one of those things is useful, then continue paying, else stop. One month's subscription is always worth it. But it seems to have given me a download, which is the "India N-CAP" something or the other. And that's a 12MB download, which is large. It contains all of the files. And I totally believe that it has covered every one of them.

**Anand**: [26:46] So, yeah, that was it! So the learning here for me is, and thanks Tanay for that input, which is you can just use the browser extension. Not the only way we could have done it.

**Manju**: [26:59] Anand... can we, you know, use the general simple mass downloader that comes with a Chrome extension which is generally free?

**Anand**: [27:09] Absolutely. Totally. And in case that fails, go to Claude because it's smarter. **It's literally like telling an assistant, "Get this done." And the assistant probably has more human-like tricks up their sleeve.**

**Rohit**: [27:24] But Anand, building... Anand, sorry, building on Manju's questions, so suppose I use this Chrome extension, it downloaded. Now, suppose I, Manju, Sajeev, whoever here, has a doubt whether it has done a job or not, **I can put the download and the link on my... on my AI and say "Check whether it has downloaded all or not." Is that... would you suggest, recommend that route or not?**

**Anand**: [27:56] Absolutely. And to be fair, for this one, it's a relatively easy check because I just need to probably count the number of files. But one another way of doing that is to open another session. Now, this is an independent agent to which we can give an independent task: uploading the entire data set, telling it, "Check what is missing." **Now that prompt is important: "Check what is missing." Because if you ask it, "Is it matching?" it might find 99% matching, "Yeah, it's matching," and it might get lazy. "Check what is missing" is an invocation to find the error, and that makes all the difference.**

**Anand**: [28:44] Or you could, instead of passing it all the files themselves, just pass the list of files that it downloaded, whatever is easier. I wouldn't worry too much about it at this stage. But the important... the other important thing to remember is that **another chat is almost like another person. So you have some degree of independence chat-to-chat. The advantage is that therefore you can do cross-checking easily, or distribution of tasks to multiple people easily.** The disadvantage is that therefore they don't have memory unless you continue the conversation in the same chat.

**Anand**: [29:26] So now that we have these files, why don't you download them and we'll start playing around with these in parallel? The files are at files.s-anand.net... I think this is the link. Yes. So I'm pinging a zip file which has all of the air quality index data. And what we are going to do is create data stories out of here. And what I mean by "we" is all of us. **Use your coding agent of choice.** ChatGPT, whatever.

---

**Anand**: [00:00] ...ChatGPT 4 or ChatGPT 4o, and Claude 3.5 Sonnet is what I'm going to use. But feel free to use Gemini or whatever. **Paid versions will give you better results**. Perplexity might not work as well. Free versions of some things might not work. And that's okay. **The point, in fact, of this session is to discover not just what is possible but what is not possible.** For which model, for which version, is it not possible, and how do we move to something else? If it's for the research, so a failure is exciting.

**Rohit**: [00:35] Anand, is it downloading for others? I can't download.

**Anand**: [00:42] No, it's not downloading for me.

**Anand**: [00:43] Okay, sorry. I'll just replace the HTTP with HTTPS. I think that may have been the issue. I've just updated the same link. Do give it a shot.

**Rohit**: [00:54] Yeah, now it is.

**Anand**: [00:56] Okay, sorry about that.

**Rohit**: [01:02] Yeah, now it is.

**Anand**: [01:03] I'll tell you how I would analyze this data. I would add this dataset, AQI.zip, and we'll do it in multiple ways. My prompt to the first assistant—and I'm feeling very lazy right now—is: **"Blow my mind with fantastic insights that nobody would have known about."** That's it. Let this guy run.

**Anand**: [01:40] And in parallel, let's ask another kind of question. Now here, I'm going to try a slightly different style. Remember we have all the answers that you had looked at? So I'm going to tell it, "I'm running a workshop. This workshop has a mixed audience, and when I asked them the question, 'What kind of datasets they want to analyze?', they gave a bunch of answers, which I've pasted here. **Now, based on this, I want you to think like a psychologist, like a journalist, like an expert who can infer human preferences and behavior from evidence. Apply that expertise to these answers and infer what kinds of insights would be really useful, interesting, and powerful for this kind of an audience. That's your first task. List those, and then find out which of those you can actually answer from this data and which of those you can't answer from this data. Pick the ones that will have the highest effect, impact, and ease. Run those analyses and tell me the most interesting things that you can find from this data."**

**Anand**: [03:04] So now, in this case, I'm taking it one step further, which is: analyze the data, but look, I'm telling you who the audience is, and I want you to use my knowledge of—or your expertise—and then answer the question. This is the second prompt.

**Anand**: [03:25] Now, I've consolidated many of these into... in the first case I just said "give me insights." In the second case, I defined how it should find the audience and then get to the insight. For any kind of analysis that I do these days, I follow the process which I've distilled over the years. And that process is in the **Data Analysis Template** that I have in a link that I will just put in the chat. Give me a few seconds for it to appear. Yeah.

**Anand**: [04:12] **This I often just copy-paste and put it into the chat after explaining what I want done. Quite often, I just upload a dataset and copy-paste the bulk of this, use it as a prompt, and have it do the analysis.**

**Anand**: [04:32] Here is the question that I will add to our survey form, which is roughly what would be the air quality index analysis. "What analysis are you or your agent planning on the air quality index dataset?" I used... hold on. And save this. This should appear as a question on the same page. Past answers should be saved. Yes. This one was... I'm not answering question seven because I didn't have a point of view. But what analysis am I or my agent planning to run?

**Anand**: [05:30] We'll find out in my case. But **if you have a prior point of view about what analysis you would like to do with that quality index, just put it in because it's a prior. And if your agent suggests something, let's go with that and see what we get.** My agents are still running. Has anyone's found anything? Just feel free to unmute yourself and share.

**Narendra**: [06:05] Yeah, I managed to create a dashboard that will visualize the AQI.

**Anand**: [06:13] Wonderful! Let's take a look at it. This will be lovely to see.

**Narendra**: [06:17] Okay, I will share my screen. Yeah, let me share. Just a minute. Hopefully, you can see my screen.

**Anand**: [06:33] It's coming up. Yes, now it's visible. Hello! Could you talk us through this?

**Narendra**: [06:47] Yeah, so basically I just gave a very simple prompt where I wanted it to give a good analysis or a beautiful visualization in a dashboard format of all the data in this AQI file. So it has created this entire thing: 267 cities tracked, so many daily records, you know, what is the monthly average, what is the trend? Is it getting better, is it getting worse? What are the seasonal patterns? Which are the worst and cleanest cities? So **just a single prompt has sort of given a very substantial visualization and useful data.**

**Anand**: [07:26] Fascinating! And Narendra... let me... and these visualizations would be in, sorry, what file type? I don't know how many of the others here... okay.

**Narendra**: [07:43] I guess Anand would know better. Basically, you can see the thing running on the side, which I'm scrolling now. It has created this HTML file basically.

**Anand**: [07:56] And I just share the HTML files on email sometimes, sometimes directly. But **the prompt that helps me in such cases is simply "Visualize this as a single-page HTML file"** in case I need it in any kind of specific format. Anyone else found anything interesting?

**Ashwin**: [08:21] I'm sorry, can I ask a question here?

**Anand**: [08:24] Please, Ashwin.

**Ashwin**: [08:25] So, you know, my focus is on energy, especially India's oil imports, especially now. So often in my stories, we need data. Now, for instance, in my prompt, I asked for a comparison in a bar graph of India's oil imports from Russia, Saudi Arabia, and USA from 2016 when Trump took office till now. So it said, "I can't give you that because you have to subscribe." But it gave me a bar graph anyway. Now, is there any tool whereby I can check the authenticity of the figures? Because then I typed in later, "What is the basis? Show source." And it just gave me a list of reports. Now, it's obviously Reuters, etc., but one can't totally rely on that. So **is there any tool by which how you can check that ChatGPT is giving you the right data?**

**Anand**: [09:31] Excellent question. Could you share your screen and I'll guide you on this? Because this brings us to the very next topic that I was going to cover, which is: it's possible to get insights, and I'll in a few minutes share the insights that my earlier prompt has shared, but how do we know if it's right or wrong? And would Ashwin want to share? No, no, please just go ahead and share your screen, Ashwin.

**Ashwin**: [09:56] Yeah. Oh, you mean how? Okay.

**Anand**: [09:59] There's a third button at the bottom, laptop up-arrow button.

**Rohit**: [10:04] Ashwin, you will have a... at the bottom of the screen, something called "Share Screen" button.

**Ashwin**: [10:11] I've pressed on the "Share Screen."

**Rohit**: [10:14] It would be doing the "Entire Screen." Yeah, it will ask you "Entire Screen" or just... so just maybe "Entire Screen" please. You have a lot of secret information on your laptop? I was just joking.

**Ashwin**: [10:24] No, no, I'm not... no, I'm just joking. I really don't know how to share screen.

**Rohit**: [10:30] You can trust everyone here.

**Ashwin**: [10:33] No, it's not that. So I've... "Anyone with this link can see this conversation" is what I'm getting. But I'll... since you can't see it, I'll tell you what it says.

**Anand**: [10:49] No, no, no, no, no. We'll do it the proper way. So now, can you switch to the window where you can see my face?

**Ashwin**: [11:02] I've seen your window differently. Yes. Okay.

**Anand**: [11:07] And you will see at the bottom a button looking like what I'm pointing to on my screen. I'm sharing my screen now. I will stop sharing in a second, but you should see, once I stop sharing, an icon that looks like that laptop on the screen. I'm now stopping sharing. Could you press that little laptop up-arrow icon at the bottom?

**Ashwin**: [11:30] I have on the top where it says "Share." I don't have at the bottom.

**Anand**: [11:36] Okay, good enough. Click on the top.

**Ashwin**: [11:38] And you're still able to see my face, right?

**Anand**: [11:42] Yeah. So it says "Anyone with this link can see." I don't know after that.

**Anand**: [11:49] Got it. Then that is still the ChatGPT window. Okay, fine. Now this is an interesting problem I'm keen on cracking. Any suggestions on how we can guide Ashwin?

**Rohit**: [12:02] Ashwin, you know, I think... are you on ChatGPT? Can you see whether on your screen URL is it saying meet.google.com or chatgpt.com? What is it saying in the URL?

**Ashwin**: [12:15] It's saying meet.google.com. But my ChatGPT, it's gone to... see, I'll tell you what there is. It's given me a very nice bar graph. It has given me what the chart shows. And then on the right-hand side, there is an icon which says "Share." Now I've clicked on "Share" and it says "Anyone with a public link can see this." So, yeah, so I don't know why it's not coming up.

**Anand**: [12:43] Got it. No, that's okay. You can share it that way then. Click on that link, please, Ashwin.

**Ram**: [12:49] Ram is saying something. Ram, you want to advise this?

**Ram**: [12:51] Yes. Are you using Windows or Mac?

**Ashwin**: [12:54] I'm using Windows.

**Ram**: [12:56] Okay, if you do an Alt+Tab...

**Ashwin**: [12:58] What I can do is I'll copy the ChatGPT link, then maybe all of you can... in the question box, then all of you can open that.

**Anand**: [13:05] Let's do that.

**Ashwin**: [13:08] Yeah.

**Anand**: [13:18] If you've pasted it in the chat window, it hasn't appeared yet.

**Tanay**: [13:25] I think other users are not allowed to send messages because I also did try about 15 minutes earlier, but it didn't go through.

**Anand**: [13:34] Okay, let me ask... let me put this in the survey form. I'll paste... Ashwin, so you all can see.

**Rohit**: [13:41] No, but Anand, Tanay is right. Your message wasn't delivered for some reason.

**Anand**: [13:47] Yes, great. So what we'll do is use the survey form. ChatGPT link and say, question: "What is the link? Shared link to your ChatGPT conversation." And this I'll request for now Ashwin just you fill this out. And I should be able to see it and I can paste it. This will be question number nine in the survey form. Are you in a position to paste it there? Okay, Nivedita is able to submit.

**Tanay**: [14:24] Oh yes, somebody's able to, but I couldn't earlier.

**Anand**: [14:30] We don't have Ashwin. Hope we didn't lose Ashwin. Well, once he joins back, we will continue.

**Anand**: [14:40] So now the insights that it's shared, and we'll pick up the same theme: in the first chat where I just said "blow my mind," it said **"Delhi is not the villain, it's the poster child. Ghaziabad was worse most of the time. Greater Noida was worse more than half the time, so was Bhiwadi."** "India has a calendar. November is not the worst; it's a significantly different regime." "Pollutant changes with the season." "**Diwali is visible, but not always where you expect that.**" That is interesting. And it compared Diwali against the next day against non-Diwali days. "Median lift is usually positive, but **Chennai's AQI rose by 61 points despite being far cleaner.**" So, yeah, Diwali just completely kills Chennai. I didn't know that. It's improved a lot lately. But okay.

**Anand**: [15:42] Well, you know, this actually, let me pause here to just reflect on what I said. I just got an insight. This is something that I did not know. It is a surprise to me. Many of you may well know, but I live in Chennai. I thought the opposite: that in fact, Chennai's rise would probably be less than other cities. So this is not only a surprise; it is a pretty useful surprise because then I can use this as a fact to let people know to tone it down during Diwali. So it's an actionable insight as well. And it goes on for a whole bunch of these. I am going to share this version of the story in the chat window. And now I just invite everyone, please just any chat that you had—Claude, ChatGPT, whatever—sorry, I say "ChatGPT," I will rephrase that to your ChatGPT, Claude, or Gemini, or whatever, any other AI conversation, and put that in here. In my case, this was mine. You may have had multiple conversations; please just share any one conversation.

**Anand**: [16:52] The more interesting conversation was this one, where I asked it to guess what the audience would be interested in. And it looked at what the audience profile was, possibly what can and cannot be answered, etc., and came up with a whole series of these answers. And this one, it says the best story to tell is: "**People think pollution means Delhi, but when we analyze this data, Delhi is only part of the story. Smaller cities...**" yeah, we already saw that. "**The real killer pattern is seasonality. Several cities become 200 points worse in winter than in monsoon.**" So it's not which city is polluted; it's which city becomes unlivable in which season, which introduces the timing angle. And that is a sharp one.

**Anand**: [17:42] Now I want to do two things. I want to cross-check this exactly as Ashwin said. Ashwin, can we try this again?

**Ashwin**: [17:56] We could, Ashwin. Have you... hold on. Yeah, I'm just trying this again.

**Anand**: [18:03] Now you could put that in in the form, in the survey. It's the last question on the survey. Once you do, please let me know.

**Ashwin**: [18:12] Yeah. Now I want to go to ChatGPT... click on your screen, click share... ah, good that you called in the experts. Yeah. And then... let me stop sharing. Okay, yes, Nivedita.

**Anand**: [18:29] Can you see mine?

**Ashwin**: [18:31] Yes. And... oh, interesting!

**Anand**: [18:36] Okay. Now, one of the things... and the reason why a shared screen is so informative is that I can infer that **this is not a separate conversation in a new chat that you're having. You're continuing a long conversation which probably includes many other things as well.**

**Ashwin**: [18:57] I'm sorry to interrupt. You see, I've been having this problem for a very long time. ChatGPT, as a journalist, is excellent for giving me statements. You know, foreign minister has said, sometimes statements of, you know, foreign delegates in foreign languages; for the Chinese, it's especially good. Where I am finding a problem is with data. Now we are very much into, and I'm sure many people are, into comparative analysis. Now **the frequent problem I have with ChatGPT and others is that it gives me data but it does not show me the source.** When I ask it for a source, like you're seeing on screen, you must be seeing below it shows Reuters, BBC, etc.

**Anand**: [19:48] Actually I can't. I see what you're seeing on the screen. Maybe you have two screens? Okay, let's look at... ha, yes, now I see it. It shows you a bunch of sources, yes.

**Ashwin**: [20:01] Yes.

**Anand**: [20:02] Now, could you click on Reuters, any one of those Reuters? And open the link which is exactly what you were having.

**Ashwin**: [20:12] Correct.

**Anand**: [20:13] **This is where it claims to have picked that particular statement from.** Correct?

**Ashwin**: [20:19] Correct.

**Anand**: [20:20] Now let's do something. Could you type in what I'm about to tell you... no, let's make it easier. I will paste something in the chat window that you could type. Or the way I'm going to do that is by dictating myself to ChatGPT and then I will paste it in the chat window. So let me begin my dictation.

**Anand**: [20:38] "**You've given me a bunch of statements. I want you to find every single error in each of those statements and list them. For any claim where you are unable to identify that error, I want you to cite verbatim from the source and include a link to the source so that I can cross-verify it. The more errors you find, the better you have achieved your objective. Make sure that even trivial errors are identified so that I can correct them.**" I have just dictated this into ChatGPT and I am going to paste this in the chat window. Could you take it from the chat and add it to your ChatGPT conversation?

**Rohit**: [21:32] Anand, also, I don't know for the benefit of other participants, could it also be a question of what version of, you know, AI you are using?

**Anand**: [21:52] See, Ashwin... sorry, I thought you were asking me. This is my paid version.

**Anand**: [21:58] Okay, let's come back, Ashwin. Yeah, this does not appear to be the Plus version. It is the Go version, which is the 400-odd rupee version, which is better than the free version. Now on the top, you have a dropdown that says "ChatGPT." Could you select that?

**Ashwin**: [22:23] Yeah, yes. Okay.

**Anand**: [22:25] This effectively tells you that there is a ChatGPT Plus that you can upgrade to, which will give you a little more control over how much thinking the model should do. **In other words, you are likely to get fewer errors in ChatGPT Plus (which is 2,000 rupees) than in the 400 or 500 [rupee version], than in the free version.**

**Anand**: [22:51] Second, this conversation continues from a dozen previous conversations. **That's roughly the equivalent of telling somebody, "Look, I want you to buy a car and I also want you to fill in my bank statement while you're at it. I also want you to plow my fields and while you're at it, I also want you to make sure that my school education planning is sorted out." And by the way, it's going to diligently get confused.** **New conversation, new topic, new chat.** This is a classic issue and that will improve the quality of your conversations.

**Anand**: [23:29] But let's go back to the chat and see what errors it has found. So that's a good tip: **you're saying that once you switch topics, you should go to a new thread.** Correct. New topic, new chat. And yes, there are both errors and sources, and it's listed a long series of these. Yes. Now, instead of bothering to read all of these, you could just add one prompt: "**Fix and rewrite**" or "**Fix and redo**." Just type that in: "Fix and redo." Okay. Enter. And you can stop sharing your screen now. Once this is done, you will get a set of results which will be better. Not perfect, but better.

**Ashwin**: [24:25] So finally, I'll just end this: **if I upgrade my ChatGPT version, will it give me better data results, or do I have to fix my prompt?**

**Anand**: [24:37] **Both.** Got it. The upgrade will help independent of fixing the prompts. Okay.

**Anand**: [24:47] Let me share my screen. That was effectively the next topic. And so what we've done so far is: ideated to find data sources, had AI download the data sources, had it identify insights, verify those, and we are next going to have it visualize them. Narendra showed us how...

**Manju**: [25:11] I've got ChatGPT to visualize it, but for some reason, I can't post it on the chat. I've downloaded it, and... okay.

**Rohit**: [25:25] Yeah, Anand, people I think only you can post for some reason.

**Manju**: [25:31] Yeah, okay. Yeah, I can't share. It just says "Your message failed" or something.

**Anand**: [25:36] You put it in the last question of the survey, Manju. That way I can present it as well for everyone and I can post it.

**Manju**: [25:45] Yeah, I can't share the screen as well.

**Anand**: [25:49] Not a problem. Can you open the survey?

**Manju**: [25:52] Yeah.

**Anand**: [25:53] Go to the last question.

**Manju**: [25:55] Okay. Yeah. One minute, why is it asking me to sign in again? It sometimes does.

**Manju**: [26:13] I've signed in. "What analysis are you or your agent planning on AQI?", the last one. "What is the shared..." yeah, yeah, "What is the shared link?" Okay, let me see if I can... yeah.

**Anand**: [26:50] All right, that has appeared. What I will do is open that link and I'll share my... Oh, fine. This is a link that I cannot see. **You have to go to the "Share" button on the top right.**

**Manju**: [27:04] Okay, right. One second.

**Anand**: [27:09] But in any case, you can't share that now. Don't worry, I will share my screen and we will talk through.

**Anand**: [27:17] So now let's say we want to visualize this. I'm going to share my screen and the specific insight that I want to visualize is this thing: which is that the question is really not which city is polluted; it is which city has become unlivable in which season and why. Now, **ChatGPT is not the best visualizer, but it's not bad.** So let's give it a shot. I'm going to ask it, "I would like you to create a slide, a single page, that visualizes the storyline that you mentioned, which is that the question is not which city is polluted, but which city becomes unlivable in which seasons and why. Make sure that you pick the visual representation which, when somebody looks at it, becomes immediately obvious as to what the key point of the story is." Actually, I could go on for a long time about how I want it visualized, but that's not the point. I am therefore going to delete this and just say, "**Visualize the key insight in a single slide.**" Let's just try that; I know it's a very vague prompt. But I am going to... oh, sorry, there's one change that I will make, though. I will ask it "**in a single page HTML slide.**"

**Anand**: [28:58] Now, what was that little bit? **If I said "slide," it might think PowerPoint slide. I don't want a PowerPoint slide. It might even have thought an image, and ChatGPT is fantastic at generating images. So, in fact, we should do that also.** But what I really want is HTML so that it can draw charts, and it's HTML as a language that it understands really well.

**Anand**: [29:28] The other thing that I'm going to do, however, is let me create a branch in a new chat. What that does is takes everything up to this point and lets me continue the conversation in parallel. In this new chat... earlier chat I was asking it to continue to create an HTML story. Out here, I'm going to ask it to **"Draw a really nice R.K. Laxman style or comic panel that communicates the main story."** Let's just see. I'm going to go for...

**Rohit**: [29:56] It's 5:30.

**Anand**: [29:59] Yeah, we are coming to a close.

---

**Anand**: [00:00] And let it draw it. And why limit ourselves? Yeah, okay, let's just do that. It can create images pretty well. If you just click on the plus icon and select "Create image," you can choose the size you want and all that. I'll let it pick automatically and let it draw. So now we have two things running in parallel. One is where it's writing some HTML and we'll talk about how exactly we can use this. In parallel, it is also creating the image.

**Anand**: [00:30] **The images can be surprisingly sophisticated.** The reason I'm saying surprisingly is we normally don't think of using images for many things. For charts, we say, "Oh, it has to be exact." For text, we say it can be just written as text for editability. But **entire slides can be created as images if you set "Create image."** And if you wanted to make a change, you just tell it to make a change, it'll make a change. Within reasonable limits, the capabilities are still improving. And what I'm finding, therefore, is many things that I would have not thought of using images for, I'm now starting to use images for.

**Anand**: [01:18] Infographics, South China Morning Post style infographics, for instance, are great examples of something that it can create if you just give it all the raw material. **Many of my talks, I compress into single-page infographics.** To show you an example, here is my talks page and this is... where is this? I had summarized this entire talk into, yeah, this comic panel. And **the comic panel explains step by step in eight panels what exactly I covered in my talk.** Now, this was one-shot, meaning I gave it the transcript of the talk and I said, "Create a comic story," and it created this. That's an entire future in itself.

**Anand**: [02:11] Or, let's take another story. This is a session that I did for National Institute of Engineering, Mysore, and if I open this image, what it's created here is a sketch note starting with section one which leads to section two which leads to section three, covering what was explained here. Again, comic style, but slightly different format. In fact, there are so many different formats that I had to use AI to give me a list of the different kinds of formats that I can use. For example, text formats. "Help me understand what are the different text formats. If I just took Rock-Paper-Scissors, what are the different ways in which I can explain it?" **Exploded diagrams, it said, is one such format, or alluvial flow diagrams is another way you can explain it. Or there is a cross-section cutaway that you can use to explain it.** I don't know what these are; it came up with the descriptions by itself. They're probably some famous representations. It created those representations from the prompt. And now I use this like a gallery, which you are welcome to take a look at. I have put it in the chat window.

**Anand**: [03:29] And therefore, **the power of images is now a little more sophisticated than it used to be.** So here is: "Delhi gets a headline, winter does the rest." Okay, it's an interesting one. Certainly something that could lead into the article. And the effort involved here, again, you saw, was very low. It's just... in fact, the bulk of the effort is thinking about the need for an alternative representation and putting it in there.

**Anand**: [04:10] This on the other hand, where it's... okay, analysis error, but that means it may have tried again. Let's keep going to the bottom. Okay, let's run a preview in full screen. Okay, so this is the HTML single slide that it created. "Delhi's headline, winter is the story." And it's showing the comparison. Okay, this is not a bad choice of chart. **AQI in Purnia, Delhi, etc., 103 to 319. That tells me the story.** I might have included the zero axis as well, but 74 to 274, 86 to 289... you may not be able to see this very well. I'll zoom in a little... okay, the zoom in doesn't work well on this. But this is enough for me to take as a single panel visual representation and put it out there.

**Anand**: [05:11] In other words, you are going to be able to put in any analysis, continue the conversation, ask it to give you a visual representation, and it'll probably do a decent job. How do you share this? You can simply download the file. It says "Downloaded the code file," and there will be a file somewhere... yeah, in this case it's called preview.html. If you open it, that opens the file in the browser. If you mail it, it'll open the file in that person's browser. Give it to somebody technical and say "Put it on some server on the website," they will put it on the server on the website.

**Neha**: [05:49] I have an interesting thing. **I had given this data to Claude, and the first point it has said is "Ghaziabad, not Delhi, was India's real pollution capital for years."** And I can't see Ghaziabad in this list of cities.

**Anand**: [06:06] Because I told it to deliver a different story. It already told me that Ghaziabad was worse than Delhi, 63% of the time. But what I asked it to do was tell me the story... the main story. And as far as it is concerned, the main story is "Monsoon vs. Winter gap." And in Ghaziabad, that gap is not as stark, maybe.

**Neha**: [06:33] Understood.

**Anand**: [06:35] And... but that is a useful comment. Because then I can say, "Hold on, Ghaziabad is not on that list. I can understand why you may not have included it—the gap between Monsoon and Winter is not as large in Ghaziabad maybe—but **since Ghaziabad has such high pollution, in fact even higher than Delhi as per your claim, let's include it because that will give people a point of reference.**" And that's all it takes.

**Anand**: [07:07] So what I do then is pass it to an editor. The editor says XYZ. I do not listen to it; I take a tape recorder and say "Speak into my tape recorder." Take it, transcribe it, give it back to ChatGPT. **Let's do less work. Let the AI do the work.** Data stories with AI is not about AI just helping us write data stories; it is about us helping AI write the data stories. Let it do the work; we do less.

**Anand**: [07:44] So, here's my invitation to you. I will ask the question: "Visualization link." And the question is: "What is the shared link to your visualization?" This will appear as question number 10 in the survey, probably the... yeah, certainly the last one.

**Narendra**: [08:09] I can share my screen and show if that... if that could be faster.

**Anand**: [08:13] Please do in five seconds. But what I would like is for you to make sure that your link begins with "/share." If you just go and copy the link from here, I can't see it. No one can, other than you. But if you go click on the share, that is something that will give you a different link and that begins with chatgpt.com/share or for Claude it's something else, for Gemini it's something else, etc. **That shared link is what I'll request you please paste in the link to your visualization.** Do give it a shot, but I'll stop sharing my screen and Narendra, yes, would love to see yours. And anyone else who's created, I would definitely invite you to show. After which I'm going to hand over to Varun for a few minutes to show how we can do some very different kind of analysis. We can see your screen, Narendra.

**Narendra**: [09:08] Okay.

**Anand**: [09:11] Can you talk us through the story? Oh, you're on mute, Narendra.

**Narendra**: [09:22] Yeah, sorry. So, rather than a story, I mean, it was like I said, it was not much of a story. I let it decide what could be the best way to show it. And I wanted to visualize the AQI data. So this is the first visualization that it did. Then I asked it to... sorry, let me start with the correct one. Yeah, this is the first one that it did. Okay. And then I went through the process of telling it to make sure that there are no mistakes and remove the mistakes, edit it, all of that. So then it gave me another visualization where it had corrected all the issues.

**Narendra**: [09:58] And then I had asked it to use this data and find other factors like temperature that might make any of these cities most unlivable and when. And that was the data prompt I gave, based on which **it says that where the air and heat both turn against you... and then it says Delhi is the most unlivable during May, while Bangalore is more livable throughout.** And then it has all this rest of the visualizations and the AQI and temperature comparison across some months that were significant. So that's how it showed.

**Anand**: [10:32] Sorry? The chart looks pretty interesting. **Could we take a closer look at that grid which was colored? And across cities, across months, what is this saying is Delhi has those maroon spots, and that's what makes it unlivable.** It's interesting that Ahmedabad also has those maroon spots. But Bangalore, if I just look at it, does not seem to be as bad. What are we missing?

**Narendra**: [11:02] Yeah, so because this is comparing temperature and AQI, and Bangalore, where I am from, I can proudly say that it is amongst the best.

**Anand**: [11:12] Oh, is that what the title was saying? That... "Most livable." Got it, got it. Fair, fair enough. And, you know, what are... and this is an invitation to everyone to pitch in. If you just take that one chart, Narendra... could you please share your screen again?

**Narendra**: [11:31] Yes. Just... let me do that. This one?

**Anand**: [11:43] Yeah. And, yes. So here is a question: **what improvements would you say we would... we should make to this particular chart?** And it's an open question to everyone.

**Ram**: [12:07] Many thoughts coming to my mind, Anand, but I'll share one or two because others should take more time than me. One thing comes to my mind is that Narendra obviously does not like heat. I'm saying it jokingly. Because it has... so there is a for a valid subjective ability, so there's a... there's a heat index and pollution index weighted into these boxes. But... but so... but but the... what is probably slightly, what shall I say? It's not misleading but slightly less clear—by the way, it's a very good... it's a very good, I like the colors and also the thought, Narendra—is that **actually during summers, the pollution level in Delhi, which is otherwise a very polluted city, is far less than what it is in winters. But since equally weighted—I'm guessing it's equally weighted—the heat and... but heat peaks.** So there would be no respite. Yeah, I'll just stop there.

**Narendra**: [13:14] Yeah, wow, that's a good insight. I mean, **if I was just to look at livability, because that's what Anand started off with, and maybe if I add traffic times to this, then I think I'll get a totally different, you know, cut of what is the best city to live in.**

**Anand**: [13:29] And it would be great, Narendra, if you could just try this. **Take that as a revision prompt and pass it. Ram has pointed...**

**Narendra**: [13:42] Now that you've started us down this rabbit hole, now I'm going to keep experimenting with more and more factors and see what happens.

**Ram**: [13:49] Now, **one thing which I find interesting, Anand, is the November is an outlier on Delhi, right?** Because... summer months, yes, absolutely, that's understandable. Why is November an outlier in Delhi? That is very, very intriguing. And yeah, from visualization standpoint, the readability, all of those can be improved without a doubt. But I'm curious to know why November is an outlier for Delhi.

**Anand**: [14:16] My quick thing is, having written—and some others will know better because we cover pollution so much in our newspaper—**November is usually the month Diwali falls.** Sometimes in October, but mostly November. And Diwali, it just goes through the roof, the pollution, for three, four days, not just one day like Chennai. In Delhi, Diwali is a longer span. **And this is also the time when Punjab, Haryana, and Western UP farm fires are peaking.** So, so yes. So the farm fire, Diwali season, holiday, it just creates a very deadly cocktail of pollution, even though heat... so that's where as Narendra said, if we were to add more like commuting time and I don't know, two, three other things, that's... it could probably show you different results.

**Ram**: [15:10] That's very insightful, right? Thank you.

**Anand**: [15:11] Yeah, low wind speed also, Tanay is saying.

**Tanay**: [15:17] I was pointing out low wind speeds, which are also a key reason at that time. Understood.

**Anand**: [15:23] Yes, of course. Yeah. Ram's question in itself is a point of insight. What I mean by that is, **as a process, that iteration is helpful.** And one of the things we can do is have AI build it, show it to people, take the feedback, literally dump it back. "You had questions that I was asked. Research their questions: why did that happen? Incorporate their improvements. Fix their confusions." And because the cost of generation now is so low, and the cost of verification is also so low, the effort can shift entirely to taking feedback. Which means that **the amount of feedback that we used to take can now be much larger and it can be crowdsourced as well.** And I mean that in a couple of ways. One, we could ask AI—another autonomous AI—"I've created this. Pretend that you are an audience. How would you read this? What are the questions that pop into your mind?" It'll give you some questions. Humans will also give you some questions. **Taking any combination of these and feeding it back means that the iteration cycle allows us to create something far richer, far more sophisticated, perhaps even a series from a story.** Yeah, Neha?

**Neha**: [16:43] Yeah, hi Anand. I have a question, something linked to the stories that we've been talking about, and I was going through the prompts that you had written in the start of the session talking about how we can have the guardrails of the kind of information that we want. But in my experience, I have found that after a while, whether it is ChatGPT or Claude, it starts to move in... in a loop. So **the stories start to become repetitive.** And because if you're someone who has been working without AI, so you know how you would want your story to move around. But in my experience, I have seen even... probably I don't know how to do it... that **even after putting guardrails of let's move into novelty, let's move into something that has not been unearthed till now in the data, after a while, maybe because the conversation becomes a little long, it... it looks like a loop.** So is there a way to for us to kind of not fall into that trap and maybe start a new chat or something? Do you have any suggestion on that?

**Anand**: [17:51] Yeah, great question and several thoughts on that. Let me list a few. The most important, maybe what you are already doing, which is: **give up and move on.** The cost is relatively low. Try something, it didn't work, it's okay. As long as you're not committing, relying on AI, it's relatively low cost. In terms of the approaches that I try, well, okay, so I have not been stuck in a loop for a very long time, maybe because some of my practices have habitualized to the point where I know how to get unstuck.

**Anand**: [18:31] And here is what I do. Number one: **try a completely different model.** Trying a new chat, I should in fact begin with, because **I never continue the same chat.** New topic for me is always a new chat. Please absolutely do that. Because the longer a conversation goes... every new chat in the conversation is like you repeating to the model the entire history. **The underlying models don't actually have memory. They replay the entire conversation once.** So if they're listening to two hours' worth of discussion, yours and theirs, and then at the end you say "Do X," it's confused at that point.

**Anand**: [19:21] So: **new chat, definitely. New model.** And new model comes in two flavors: either a more advanced model within Claude or ChatGPT or Gemini—just keep selecting the top dots, for which you may need the paid version—it's absolutely worth buying the paid version, there is no better thing that you can do. **Or switching over to a different provider, like from Claude to Gemini, Gemini to ChatGPT, whatever, because each has its strengths and weaknesses, whatever.**

**Anand**: [19:53] On top of that, improving the prompt can help. But the way I improve the prompt is this: **I copy the entire history of the conversation, give it to a model and say, "Look, take a look at this. This is utterly repetitive. What should I do to make it more novel? You give me the prompt."** And this is meta-prompting. You use AI to teach you AI. Usually, I suspect you will find far more powerful improvements even before you get there. Narendra, could you stop sharing your screen, please?

**Narendra**: [20:29] Yes. I just had a question, Anand. Just... see, this is on the... and it's slightly philosophical, so please bear with me. So this is based on what you said about feedback—that you keep using its own answers to give it feedback and sort of not fall in the loop, but actually refine and get a better answer, better answer. Now my question is, if you look at the observations that people had immediately—somebody talked about how November is an outlier, or how winter... may the AQI is higher but then the temperature is sort of creating a spoilsport—now, these people within one glance had created some kind of an instinct or understanding of the data. Now, **if I stop trying to internalize what I'm seeing and if I start sort of recursively trying to get a better and clearer answer, after the third or fourth prompt, I will get the perfect answer, but I might not have really internalized or learned anything from that exercise.** And I'm just wondering whether that is a pitfall we should worry about.

**Anand**: [21:34] A very good point. How do we judge somebody or something which may know more than we do? To which my thought is: isn't that what every editor does? Isn't that what every judge does? Isn't that what every auditor does? Arguably, isn't that what every teacher does? Isn't that what every coach does? We are guiding intelligences that are more advanced, and our ability to steer them, our ability to verify, is not necessarily constrained by our ability to reproduce what they do, understand fully what they do. We have evolved several mechanisms. And this whole notion of verification, steering, improving smarter intelligences is such a huge topic that I will just share one prompt suggestion. "**How do professions guide or steer smarter intelligences or smarter people than themselves when they don't have enough information?**" Ask this to any good model and it'll give you a series of thoughts, thought-provoking ideas that you might have.

**Anand**: [22:46] But with that, I'm going to do two things as we wind down this session. First is, Varun, could you share your screen and walk us through the geo-spatial analysis that you've been working on? So far we've been looking at... and I'll be talking through it while you share your screen... the initial part that we did all based on structured data, mostly. some amount of textual data, but it was mostly numbers. Can we visualize geography? What Varun's done here is... and now over to you Varun, maybe you can just talk through what you've done.

**Varun**: [23:24] Sure. So like, I have used LLM to generate a geo-spatial insight over any area. So what I did, simply, I created a dashboard using the help of some of the coding agents and later on I... I moved to the manual part where I just see any of the large patches over the area and see, zooming over into that, like and see if any... change what are the... location looks like and rather than the research by myself, I had a... go back, go back. I'll need to explain that a little more. Please set the opacity to maybe about, yeah, 100%. So this is... which city? Sorry, I can't see that.

**Anand**: [24:09] One is Kanpur.

**Varun**: [24:12] **Kanpur. Okay. And vegetation index has fine... so one of the data sets that we have available is satellite imagery.** It turns out that we have models, OlmoEarth being one, which can analyze satellite data very easily... but I mean with moderate ease, it takes a few hours to do this kind of an analysis. What Varun's done is taken small grids, maybe 100-meter by 100-meter grids, and for each one of those asked it, like asking ChatGPT sort of but in a slightly different way, "**What is the similarity of this patch to the word 'vegetation'? Or to the word 'water'? Or to the word 'urban'?"** and so on. What you see here is the similarity to the word "vegetation." **Red means low similarity to vegetation; green means high similarity to vegetation.** And you can see that where the river is snaking through Kanpur, there is clearly high similarity to vegetation. And the red spots... Varun, can you reduce the opacity? You can see what some of the red spots are. The red spots seem a little more industrialized areas. Could you now find the darkest red spot and zoom in?

**Anand**: [25:42] Specifically what we're looking for is—and I'll make a correction to what I said—**what Varun's done is not identify what area has high red or high vegetation or low vegetation, but the change from 2015 to 2025.** So this spot has had a huge drop. Let's see what it looked like in 2015. Varun, could you change the timer to 2015? And give it a few seconds. **A lot more, well I'm guessing algae that was present at this point.** And now let's move it to 2025. **The algae has removed... has gone away, the lake has been cleaned, and there is less vegetation.** The thing that I'm learning from this therefore is that when someone blindly uses something like a vegetation index, it also reports these kinds of false positives. **A lake cleanup, which you would say "Yeah, okay, technically you can argue that that's less vegetation because the algae's gone away," is not really what we think of when we think about less vegetation.**

**Anand**: [26:57] Now, let's zoom out again, Varun, and take a look at some other red spot. Yeah, take any red grid, big red grid. And yeah, now zoom out, show 2015, please. So something seems to have come up here as development. In 2015, the picture was like this. But in 2025, the picture is like what we saw earlier. Now, this means that **we can very easily visually discover geographic stories with this kind of data.** And then write about it. So which begs a question: okay, how do we go about writing about stories like this? Can we use AI's help there also? Varun, back to you.

**Varun**: [27:49] Yes, so like what I did, I have used Gemini for this, since it is having a good visual skills and also it has a... has a inbuilt maps tool so that it can go through the location coordinates and check for that. So firstly I've just told it what it was supposed to do and how I... what my further inputs would be looking like. It will simply containing a past image and the current image and also a screenshot of the dashboard which is completely optional. It can also see through the changes directly through the screenshots.

**Varun**: [28:23] And from the next prompt, I just have to provide the screenshot of the past and the current and along with the screenshot of the dashboard how it looks like. And also I have to provide the metric which specifically I am looking for and also the coordinates of the location so that **it can really go and use its maps tool or the skill and to get know about the location, how exactly the location is, and like further search on the internet for the story which I have already asked it to do.**

**Varun**: [28:57] Like its task is to fetch the details of that location using that location coordinates or the name and also search the internet for the particular cause of those changes and verify it not from the... like from multiple sources. And if those... the reason or the cause is really a newsworthy, then **it has to write a four to five lines of article with appropriate titles in a specific style which for my case was a TOI [Times of India] reporter.** And also I have given it a power to like not only always give me a "yes" on my points, but also it should reject my request if the... the location does not is having a newsworthy story, we can say. Like it can be a false positive by the data or it might be a natural change or a normal weather change due to which it has the change has occurred. So I have given it that power also, like "You are allowed, you can just reject my request if it does not have any newsworthy story."

---

**Varun**: [00:00] ...have any interesting story. So it has further, like, gone through and this is how you are seeing, like, for the Bangalore... for the lake in Bangalore, it has generated a story, and same thing goes on for the further locations also.

**Anand**: [00:16] And the important thing here is—and I'll say maybe probably three important things here. One: that **we can create stories out of visual data as well.** Just take a screenshot, send it; it can analyze. Second: **it can fact-check or ground the stories against real data.** In this particular case—so if you could scroll back to the earlier story on Bellandur Lake, Varun—what's interestingly and kind of similar to what we found earlier, he was saying: just because there is less water surface doesn't mean that there is a climate casualty. Just reading from the first line: "This is a deliberate approach, rejuvenation, etc." And the fact that it is able to get this data by researching means that **even a false positive kind of a data-driven story can be converted into a different perspective.** The story, if I had not even done the research, would have been: "Bellandur Lake has less water." But with the research, the story now remains, but just differently, which is: "**Less water does not mean bad always. Bellandur's rejuvenation is a good example.**"

**Anand**: [01:34] The third thing that is... that we should probably think about and take away is that **Varun is not a journalist, does not have experience writing data stories, and has merely been iterating based on feedback on what kinds of stories can be written.** And the amount of... the number of iterations that it took to get to this kind of a prompt was probably about two or three iterations, no more than that. Varun, could you share this prompt on the chat? And if you don't have permission, then send it to me on the Straive chat; I will paste it here.

**Varun**: [02:08] Okay, I'll send it.

**Anand**: [02:11] Thank you. You can stop sharing now. Thank you. We have less than a minute left. I'm going to share my screen and wrap this up; we'll probably go a couple of minutes over. But the primary thing that I wanted everyone to take away—and I was also going to ask you to fill one more sheet... the primary thing is that **AI can do things that we probably haven't tried before, or have tried and failed but now it has new capabilities.** And therefore, if today—either using some of the stuff that you saw in the session or some of the ideas that it sparked in your head—you are able to craft a new data story, doesn't matter if it's useful, if it's helpful, you would have improved your skill one notch. So **one takeaway that I would like for you to leave with is something that you will try on Monday.**

**Anand**: [03:11] Please just fill in the last two questions, questions 11 and 12, and that will serve as a reminder. I will also follow up on email dropping you a note reminding you. But **what's something that you changed your mind on during this session?** Something you thought at the beginning, now you think differently. Just drop in a sentence or two. And **what's one thing you are going to try, or get somebody to try?** However, but some action. This is something that will be helpful for all of us. I'm going to create a data story out of the data that we've gathered in this session—your responses, our discussions, etc.—and will share that on email. If you have any follow-up questions, things to experiment, try, whatever, I'll be sending you an email by tomorrow. You are more than welcome to just drop me an email; we can connect and learn from each other. Any last questions for me?

**Tanay**: [04:14] Anand, not a last question. I mean, I just... one thing, just to maybe we can also think through or should do, what I did. So for this AQI thing, so I also asked, like, you know, to list down the... **can we connect the list of constructions that's going on around the metro cities? Right? That could also lead to this AQI index.** That is a major part that, you know, we could also think through. So just giving some hints so people can also try these things as well.

**Anand**: [04:46] Fascinating idea. It **combines the geo-spatial analysis that we were doing with the other data set.** Absolutely great idea. Yeah.

**Ram**: [04:54] Anand, in between I got lost actually, I mean I had... I got dropped out because of the bad internet. Can you please paste the form's link again?

**Anand**: [05:07] Certainly, I will put it in the chat. Thank you for asking.

**Ram**: [05:22] Thank you.

**Anand**: [05:24] Yeah. Just a quick check... see if anyone's updated... yeah.

**Neha**: [05:43] I had a question. So **at what stage of this workflow do you think, you know, the human in the loop... like, do we come, do we check the analysis or can we be very blind and trust that AI would have done the analysis properly, hasn't missed anything?** Or do we have to go back and check the individual numbers, you know, in terms of... this is just numbers. Like, I'm sure the analysis is right. But do you think one has to check? I mean, I check it all the time. I don't know if my question's clear.

**Anand**: [06:08] It is. And **I would treat it the same way as I would treat a fresh journalist who's joined my organization.** At first, I will check very carefully. How good is this person? Should I even retain them, should I fire them? Over time, a certain amount of confidence builds. Then I check for what's important; I don't check for what's less important. Actually, even otherwise, I may not check for what's less important. So **stratification matters.** If I'm, say, checking a medical report, I will double-check, triple-check. If I'm asking where should I... what should I eat next, give me some suggestions, right or wrong who cares? "This has more calories versus this has less calories," okay, if it's off by a little bit, I'm not going to fast. So prioritize. Build confidence. **Over time, when you have a reasonable amount of confidence, delegate the verification. Have another chat verify.** When that builds even further confidence, you have built an intuition on: this fellow tends to get these kinds of things right, I won't check these; these kinds of things I still want to check. Over time, as the models keep getting better and better, and we keep getting better at prompting, it's hard to tell what's happening. You will change what you verify, what you don't verify, but periodically go back and make sure that nothing has regressed. Long way of saying: **work with it just like you would work with a person.**

**Rohit**: [07:39] Anand, one related to this, one tip may help some participants. Like, I saw questions about which version—Go, Plus, Pro, similarly. So, you know, many of us have our own or office-given high paid version. But suppose others who don't have, there's no reason, as you said, for them not to use enough. **Is how much do you think is a function of usage?** Say, I'm a... only a Go user or I'm a free user, but I go to AI only twice a week. Now, therefore, it will... will it still give me good answers because I'm a low user, I'm not exhausting my limits and whatever use I do I can get? Or there is no such rule or there's no such thumb rule to go by?

**Anand**: [08:27] No, it's a good question. **I personally believe that this is the highest ROI investment that one can purchase.** I'll give you my rule of thumb for technology subscriptions in general. Taking YouTube Premium as an example: I've tried and dropped off YouTube Premium five times. I tried it, I used it for a couple of months, and then my usage fell. Then I stopped. And everybody was saying YouTube Premium—even my research indicates YouTube Premium is a great idea—I've subscribed to it again, but then immediately dropped off because I was not using it. Third time I tried it, I used it for almost six months, then it changed. That's okay. The only thing is that **I would, in the case of AI, reduce my period of re-evaluation. Not longer than six months. Because the speed at which things are changing, you want to check if there's something that you have missed.** And here's how I suggest doing it. **Start with the Plus version for one month. That is the 2,000 rupees.** Whether it's Claude or ChatGPT, start with these two for now. And use it as much as you can for one month. You're paying 2,000 rupees; you may as well make the best use of it. It's a fantastic psychological hack. It serves two purposes: A, you are utilizing your subscription; B, you're learning to use more AI, which is great. At the end of one month, if you're utterly, totally convinced that it's *paisa vasool* [worth the money], continue. If you're not, unsubscribe; six months later, for one [month].

**Rohit**: [09:59] Good one. That's a good point.

**Anand**: [10:04] I do invite you to fill the takeaway. The whole point of the workshop is to have a concrete takeaway. It may not be obvious to us what we should take away, but **your takeaway could just be, "I will go to ChatGPT and ask it what my takeaway should be."** Even that is okay, but have a takeaway. Ashwin, yes.

**Ashwin**: [10:24] Thank you, Anand. Rohit, I thought there was a slightly nuanced [angle] to your question, right Rohit? And that is: **does the ChatGPT or any AI service provider have an affinity to value frequent users versus infrequent users?** I thought that was also a hidden, implied question in yours. Is that right, Rohit?

**Rohit**: [10:48] Yes, yes. Because many of us are very heavy users, so we have Pro accounts. But those who may not... yes, it was implied. True.

**Anand**: [10:57] Interesting. **Frequency of usage tends to build context.** And to that extent, I would say yes, there is a benefit. Affinity? I'm not sure.

**Ashwin**: [11:13] But that's a great question, though. Anand, would love to hear your views when you research that more, right? **If the entities are sort of motivating or sort of incentivizing usage by giving higher, better value results, that would be a great insight to know.** If you do research that, right Anand, would love to know outcomes on that.

**Anand**: [11:41] **A lot of the incentivizing right now is happening on the coding agent side, not as much on the direct chat side.** That ended last year. So if you are using coding agents, yes, it's good to keep an eye out for the current monthly offers, so to speak. And there are several of those that are happening.

**Rohit**: [12:02] Yeah, fantastic. Thank you, Anand.

**Anand**: [12:04] Any more takeaways? We're kind of still stuck at nine takeaways. We can just pause for a minute, think about what we're going to do on Monday.

**Narendra**: [12:23] Anand, I'm still typing out mine, I'll just do that. Anand, before that I had one question: **how do the Chinese LLMs measure up?** You know, because everyone's using Chat in this one. But there's Qwen, there's DeepSeek, and I found that DeepSeek and Qwen are... they're pretty adept at a lot of tasks. Now, what's been your experience with them?

**Anand**: [12:42] **I use the link that I've just pasted in the chat window on LLM Pricing as my evaluation.** It's a comparison of quality on the Y-axis versus cost on the X-axis. Let me give you the quick answer: **Chinese models are good if you're cost-conscious and otherwise, highly cost-conscious at scale, not otherwise.** But here is the detailed answer. The Y-axis ranges from a high school freshman to a college junior to a Master's, PhD, tenured professor. And we had, in March '23, a few models which were a high school student level of intelligence. That grew. The big jump was when GPT-4 became as smart as a college student. And another big jump happened when recently ChatGPT-4o became—or 4.5—became as good as a PhD student. And then we have now Gemini 1.5 Pro as smart as a tenured professor. We have even better models, and this is updated as of... well... a few weeks ago, so this is already very outdated. And Fabel is not even on my list, etc. But therefore, from an intelligence perspective, they're pretty high. And **where are the Chinese models? Let's take DeepSeek. Pretty good, better than a PhD candidate.** When DeepSeek... you can see them all on the left side. But the frontier models are a little ahead. A rough rule of thumb is **they are about six months or less ahead compared to the Chinese models.** The cost is the other equation, and these are significantly cheaper and available for hosting open source for those who have a very powerful [machine]. So one way to think about it is that these models... even if all advancement stopped, then we still have a reasonably good range of intelligence-cost trade-offs. Unfortunately or fortunately, it's not stopping; it's, if anything, accelerating. So one of the effects of that is that there will be smarter and smarter models. There will be things that you couldn't even dream of doing that we will start doing. We will hunger for that. There'll be enough marketing around it that we will want to do it. And therefore, the Chinese models will have a lot of catching up to do.

**Narendra**: [15:13] A lot of catching up. Got it. Thanks, thanks Anand.

**Anand**: [15:19] Okay, we've gotten to 10 takeaways. I do invite everyone to continue filling in the rest. But let's wrap up the session. Thank you for joining in. Please do try stuff out and all the best with your data storytelling.

**Rohit**: [15:35] Thank you Anand. Thank you everybody.

**Tanay**: [15:37] Thanks Anand. Bye bye.

**Neha**: [15:38] Thank you Anand.

**Ashwin**: [15:40] Thanks Rohit. Thank you Anand.

**Ram**: [15:41] Thank you.

**Anand**: [15:42] Hey thanks, Anand. Thanks so much. Thank you. Bye.

**Varun**: [15:47] Thank you, Anand.
