DEMO | Coding Copilot Connectors to Bring External Data into M365 Copilot

Want to connect Copilot to your business data but there is no built-in connector? In this video, Microsoft Copilot DevX PMM Sébastien Levert shows us how to build Copilot Connectors using code so you can bring ALL your data in the Copilot Semantic Index.

Join me and Seb as he opens up Visual Studio and takes Microsoft Copilot to ANOTHER level! Are you wondering if there’s a way to do it WITHOUT CODE? Good news. Yes, and you can find all the details here.

Watch my 90+ courses on Pluralsight (opens in a new tab)

Video Summary

  • Connecting External Data: Demonstrates how to link external data sources, like GitHub, to Microsoft 365 Copilot using custom code for seamless integration.
  • Creating a Connection: Explains the steps to create a connection, define a schema, and set search settings to ensure Copilot understands and uses the data effectively.
  • Ingesting Data: Covers the process of fetching items from the API, transforming them into a suitable format, and loading them into Microsoft Graph.
  • Using the Data: This shows how to create an agent in Copilot that helps users identify and prioritize issues using natural language queries.
  • Practical Example: Provides a practical example where the agent helps prioritize authentication issues, demonstrating the value of integrating external data with Copilot for better decision-making.

For more information, read the transcript blog below, or watch the video above!

Video Transcript

The best way to add more value to your AI implementation is to have it trained on all your data, no matter where that data lives. While Microsoft provides many pre-built connectors to bring the data into the graph and the Copilot index, sometimes we have to code our own. This is exactly what we’ll cover in this video. Welcome back to The Ultimate Guide to Customize Microsoft 365 Copilot series. I’m joined again by Sebastian Levert, Principal PM Manager at Microsoft and an amazing developer. You finally get the chance to open Visual Studio in this one. How are you doing, Seb? I’m doing amazing. We’re going to crack open Visual Studio today. About time! Yes, I know you can’t wait for it.

Okay, Seb, so far in this series, we covered quite a lot. We started by covering the terms and concepts. Then I showed you how you can do SharePoint agents without any single line of code. Then you showed us how we can use the Agent Builder to do even more than just SharePoint without code. Yep. Then we learned how to use the built-in graph connectors to bring in a website, a file share, an SQL database, and it. What are we going to learn in this episode? So today, we’re going to learn about how we can use code to connect to the graph and ingest data that are living outside of your organization.

In the last episode, we talked about an example of what we can do with file share and an example of how we can do it with a website. What if your data is not available as one of the data sources we offered by default, but your business is still driven by that data? A good example of that type of data, and that’s exactly how my team operates, is everything is driven by GitHub. GitHub is where we store all of our user stories, all of our issues, all of our tasks, and that’s how we drive engineering to build the products that we use to ship the Copilot experience as a whole. Well, what we’re going to do today is bring the data from GitHub inside Copilot, and we’re going to be able to ask natural language questions on what I should be doing tomorrow based on my current priorities straight into GitHub.

So, is GitHub Copilot to create Copilot? We do use GitHub Copilot to create Microsoft 365 Copilot, yes. Copilot writes its own code. It’s kind of a bit of an Inception moment. I agree with you. So, if you don’t mind, folks, let’s just go through a little bit of a reminder here. Remember, M365 Copilot provides the UX, the orchestration, and the foundation model, and today our focus will be using Visual Studio Code to bring new knowledge. That knowledge is the data that lives inside GitHub, and specifically the issues of very specific repos that I care about. So, that’s it. No more slides. No more slides. Let’s go to Visual Studio. Let’s dive straight into Visual Studio.

So, I’m here in Copilot, and I need to bring data. As of now, there’s no data that I can actually ask questions on. If I was going to ask what is the next issue I should be working on, it has no idea. It’s just going to either generate something that is funny or it’s just going to say, “I don’t have context for that,” and give me some hints. It’s going to look at your emails and say, “Well, you have many issues based on your emails. You should probably work on yourself.” Very possible. But what you want to do here is to bring content directly from GitHub. So, let me show you here. I have a repo. It’s a repo that lives publicly. You can actually go see that repo. There’s no code in there. It’s just a repo I’m using for the purpose of this demo where I have 10 issues. So, I have an issue around cleaning up code for an ESPC demo and then a bunch of other issues in there that are the issues with this fictitious app that we’re building. So, there are people assigned to it. There are a couple of things that are really interesting in there. So, I want to bring that straight into M365 via graph connectors.

Awesome. Let’s do it. How do I do that? As go here. So, all the code that I’m going to show today is available publicly and freely on GitHub. You’re going to be able to use it and do the exact same thing inside your own organization. Luckily, you’re using GitHub, or if you’re not, you’re going to be able to adjust it to connect to another API and bring similar information from that API straight into your app. But all the code you do today, we’ll share it. We’ll put it somewhere public, and it will be linked in the description below. So, you don’t need to start from scratch. Focus on the video, and you’ll be able to copy-paste everything like every good developer does. Exactly. Copy-paste-driven development is who I am.

So, I’m going to start here. I have an app here. The app is built as a TypeScript Azure function. So, it’s a function that will be invoking itself on a timer. So, every couple of hours or every couple of minutes, depending on your scenario, this will wake up, do something, go to sleep, wake up, do something, go to sleep. Cool. It’s a great scenario for indexing content. Maybe every five minutes, I’m going to go get the latest, grab that data, ingest it, grab the latest, ingest the data, and that’s it. Awesome. So, we’re going to start here. I have an app that has a couple of functions that are available. Let me go to the bottom here. I have three different apps and three different functions in my Azure function. The first one is a timer called deploy connection, a timer that is called full crawl, and a timer that runs incremental crawl. So, think about it in three different scenarios.

First, you want to deploy your connection. You want to, you remember earlier when we did the setup? Yeah. We want to do the exact same thing, but only once. Okay. And then it will just recur on its own. We don’t have to manage anything from the admin center. Exactly. We won’t go to the admin center. We won’t touch any of that. We will do the deployment fully programmatically. So, the first time you run the function, it’s going to go up, it’s going to call the service, it’s going to do a bunch of things. Then afterwards, it’s going to invoke the second function, which is the full crawl. So, the first time it’s going to bring all the data, and after that, it’ll just be every minute, just bring in what’s new. Exactly. We’re going to, as part of the code, we’re actually stamping the last time we did the call to the API. The reason why we want to do this is to ensure that the next time we call, we don’t recrawl everything, especially in large data sets. We want to limit the number of data that we want, especially if you’re working with SaaS software that is making available their APIs. You don’t want to get throttled. You don’t want to get 429s. You want to make sure that everything you get is as tiny as possible.

So, let’s start with the deploy connection and see how that works. So, I’m going to go to deploy the connection here, and basically, it does four things. The first thing it does is ensure the connection. If the connection does not already exist, it’s going to go and it’s going to create it. Then afterwards, once the connection is created, it’s going to enter the schema. What is the schema? The schema is the shape of the data that you want to ingest. This is fully customizable, and it’s up to you to define how you want to ingest the data. So, you want to have maybe an ID field, a title field, a description field. You want to have an assigned to, the ID of a repo, who’s the owner of the issue, and so on and so forth. This will really depend on each company, right? And I mean, even for each company, for each connection, the single connection you do, it will be different. Exactly. Exactly.

Are there any types of fields that Copilot does not recognize or any types of fields to be careful of? I think we need to be very careful with the data that you want to bring in. You want to bring in data that is easily semantically accessible. So, don’t bring JSON payloads, for example, because then it’s going to be harder. Use JSON as the way to transfer the data, but get a specific field, a date, a string, or a number. There are a couple of these. All of those are documented on our Dev Center, which type of fields, but really keep within these boundaries. Don’t go wild and bring… So, sometimes simpler is better. Absolutely. Also, something very critical is even though you want to bring so many things, when you think about a connection in a schema, make sure you think about the simplicity of that schema. Don’t bring 100 fields. Yes, sometimes it’s going to be useful, but you need to understand the business case. You need to understand how devs and users will be using that data. If they’re never going to use one of the fields as a way to ask a question or as a way to render the data, leave it out. Fewer data, fewer hallucinations, less making things up, and better accuracy for Copilot. It goes back to that same thing every, not only for Copilot but for every AI system. Train it on the least amount of good-quality data that you need.

“Absolutely any you add that’s not needed will only confuse the model in the back. Exactly, then afterwards we’re going to go and we’re going to set the search settings. Graph connectors also have a search component, and the reason why you want to set the search settings is to make sure that we have an Adaptive card that represents the information that Copilot will use to ground itself in that content. What are the available pieces? What are the available properties that it can use to return the data to Copilot so Copilot can render the data? Okay, and finally, we’re going to start a full crawl. So we’re going to go a little bit deeper into the entering to connection. The entering to connection is actually simple. First, it’s going to get the connection. It’s going to go and say, ‘Hey, I want to get the connection. Is it already available?’ And if it’s not available, and you’re going to see it here, it’s going to fail. It’s going to return me either a 404. 404 means does not exist, not found, not found. So you need to create one or you might need to do some consent. So you’re going to, the first time that you hit the endpoint, you’re going to be prompted to do an admin consent on the APIs. So you’re going to have all of that in the terminal directly for you, where you’re going to be able to do that. And so you might need a certain admin role in order to consent to this, at least a cloud application administrator role. And then I think the piece that is the one that is the most critical to us is the create connection. So let’s go and see how we really create a connection.

Well, if you’ve ever worked in the M365 space and you coded there, there will be nothing special for you here. This thing here is a simple call to Microsoft Graph. Microsoft Graph on the external SL connections, and you’re going to pass in an ID for your connection, a name, and a description. This description is what, in combination with the name, will be used to help Copilot understand what this data is about. So when you ask questions, it’s going to automatically semantically connect the dots between what you’re asking and what Copilot is. This is what we did in the previous video when we went in and we said, ‘Hey, this is the Microsoft 365 IT Pro knowledge base.’ Yes, that’s exactly what you’re doing here with code. Exactly. Awesome. And then afterwards, we said that we were going to set some search settings, and we’re doing the exact same thing here. Again, we’re using a URL straight to the graph to the newly created connection ID, where we set, ‘Here are the result templates that I want to use.’ Okay, and that result template is a template.json. So we’re automatically using that file here, template.json. It’s an Adaptive card, and all of that will give you all the information that is available.

So it’s going to know about the owner, the repo, the title, the URL, the last modified, the author, the abstract, and here, as in, is visible to false, but it’s a hint for Copilot to know, ‘By the way, I also care about issue number assigned to innate. I don’t want to show it, but I care about it.’ So when you return the data, it’s going to automatically use these signals to say, ‘If the user wants to render something in Copilot, say you want to render a table with ID, a title, and assigned to, well, I need to have these things in here because if not, Copilot will not know that.’ So it’s a little bit of a hint here to Copilot saying, ‘By the way, it’s not shown, but I still care about these things.’ So this is still, even if we’re talking about AI and Copilot, you still need to tell it exactly what you want. Yes, you cannot assume that Copilot will just know everything and automatically read your brain, at least for now. Exactly. Who knows about the future? You’re absolutely right there. So let’s go back to my connections here. Actually, we did the ENT schema res search settings. Now we’re ready to start ingesting content. And the ingestion of the content is basically a mix of three things.

The first thing, if we go to the full crawl, the first thing here will be to ingest the content. It’s as simple as that to ingest the content in there. And here, there’s a little bit of, you can go through the comments here, but depending on, ‘Hey, if you already have a previous one, start from the previous,’ so on and so forth. Remember when I said you were the most amazing developer? You’re like the only one who actually comments on their code. Oh yeah, well, all the time. I’m actually, because it’s for myself later, it’s for the next time. So if we go see the ingest content method, you’re going to see here, we do basically three things. We get all the items from somewhere. Then afterwards, we’re transforming the items. We might be doing, ‘Okay, I’m getting data from an API. Maybe I want to reshape it a little bit. Maybe I want to call other APIs also as part of it. Maybe I want to concatenate two fields together because that’s how I want to index it.’ And then afterwards, I’m going to load it. I’m going to ingest it straight into the graph. So I’m going to go to the get all items here.

You’re going to see that it’s actually just going to call one thing, get all from API. So how does it do it? Well, it’s very, it’s kind of simple. What it does here, it gets paginated issues. So here in that case, that’s where your code goes. That’s where, as a dev, you bring all the value to Copilot. You know how to call your API. You know how to call your content. You know exactly how these things work. In this case, I know that GitHub will automatically paginate its content. Okay, I don’t want to only get the first page because if I call the API only one time, it’s going to only give me the first page. I want to be able to bring all the pages, which by default is a 100-item page. But I tell you your code never has problems. What? Your code never has problems. No, never. So you never need more than one page for your own code. You’re doing this for other people. Yes, to help you, I hope. So we’re going to go in the real code isn’t here. So basically what it does, is it calls the API. If you’ve ever called the API to GitHub, it’s this here, api.github.com/repos. We pass in the repo. The repo is an environment file, an environment variable.

Basically, you say, ‘I care about this specific repo.’ And then it’s going to do a bunch of logic based on which page I’m on, what is the last time that I called in to give me all the results, and so on and so forth. I’m building a header here to make sure I’m authenticated. And then afterwards, if there are results coming back, I see if there’s a link. If there’s a link, it tells me there’s another page, and I just basically iterate through that. Awesome. In the end, what’s going to happen is once we’re done, I’m going to create a map. I’m going to create this object here, and this is basically where we say, ‘This is the data I care about. I want an ID, and this is how you get the ID from the object. I want an issue number. I want an owner. I want a repo. I want the assigned to.’ And here you see that there is some specific logic in there. When there are multiple assigned to, you go and add everybody with a comma between each other. When there’s a last modification, I want it to be already preformatted in the way that the content is expected in Copilot.

Here, is the content, I want the title and the body of the issue. So I concatenate these two things here. Awesome. This is all business logic. In this case, it could be remapped to a Jira scenario, to an ADO, or any other platform, Bitbucket or what. But most likely, every dev will have to do this part on their own for their data. This is good for GitHub issues, but that’s about it. Exactly. And as part of the template, you’re going to see that’s why we have a custom folder. The rest can be reused almost as is. But this one is where there are things happening there. So the first thing we do is we do that. We do get the, if we go back to the ingest here, we get all the items. Then afterwards, we’re going to transform all of the items. In this case, what we’re doing here is we’re getting an external item from an item in the graph connector. An item is called an external item. So we map the content from the item coming from GitHub to an actual external item for a graph connector. So let’s go see what we do here. We basically do one thing here. We return for each object an ID and a set of properties. And these are the custom properties. So we have our schema that was ingested, and now we map every property from the schema straight to here. The one thing that I want to bring your attention to here is the last one, ACL, Access Control List. In this scenario, we make it available to everybody. So it’s a public repo. We want to make it public. So if I go here, get EF from the item, you’re going to see it’s a very simple ACL schema.“

“We grant everyone access to every single item, but this is where, as a developer, you can actually bring in who has access to what. You can map it to Azure Active Directory identities, to a user, to a group—you name it. This is really important because you don’t want to bring in HR or financial data and have it available to everyone. Exactly, and it’s critical for everybody to do that. So, in this case, in this scenario, it’s a simplistic scenario, but at least we removed the ACL from this actual map so that way you can bring, for each item, that as part of it. Then afterwards, if we go back to the ingest part of it, which I think is at the bottom—yeah, at the bottom here—we do load the content. Now, load content is the following: again, we’re doing an API call to a specific endpoint, which in that case is part of the connector ID we created earlier. Then we’re going to do a PUT to that item. A PUT is a way we can call it also upset, so it’s either a create or an update on this item. Now we pass the entire item, the item we just created, the full with the ID and the set of properties in the ACL we put into the system automatically. Now the item is available on Graph, and very shortly after, seconds after, it is now available as part of Copilot to start searching on. This basically gets called every minute through the UAL function that we had, exactly, either through the full crawl or through the incremental call that does exactly the same thing. Awesome. So, depending on where you sit, you’re going to be able to do that.

So, if I run it, and actually let me do that, I’m going to go and delete a temp file. So here, restoring as part of a temp file, the last crawl—when was the last crawl? So, just for the purpose of this scenario, I’m going to delete the entire folder. I’m going to run this app here, and it runs straight in here. It runs using the local Azure function runtime. You can also host it directly on Azure already; it’s not an issue, it’s all there. And now it’s going to go through all of the steps that we just talked about. So, it’s going to start by running the deploy connection, saying it’s the first time I run, I don’t know exactly. Then it’s going to tell me, “Oh no, it already exists, your stuff already exists.” And then now it just went into—it went really quick—it just starts a full crawl. And why does it start a full crawl there’s no file that says last crawl, so I need to start with a full crawl. And it’s going to put this item right here. Now it gives you all the details of all of that, and now it just goes on and on and on and on and on. So, it gives you a log, and I only have 11 issues right now. Good. And it just started the incremental crawl, and it ended the incremental crawl. It’s not even a second, it’s not even a second, it’s not even 100 milliseconds. And why did it do that? Well, we just saved back the incremental. But let’s say we go here and say, “We have a bug in the auth module for the mobile app.” Okay, I’m going to save that. Now we’re just going to wait for a couple of seconds, and now it’s going to rerun this function. When it’s going to rerun the function, it’s going to highlight that, “Oh, wait a second, I was able to call the last time at this specific date.

Now, is there anything that was updated since that date?” And it’s going to say, “Give me a second, I’m going to call the API.” The API will be called, we’ll do the fetch, we’ll return back to our full pipeline, we’ll go through the ACL, we’ll go through mapping everything, and now it just updated automatically. Now I have, wow, a title, “Bug in the mobile app.” But that’s really also dependent on your API. So, if your API only looks at the issue ID, for example, this would not have worked. Exactly, you need good quality APIs to have a good quality experience. Exactly. You have two options here: either you build your API and maybe even build your own ingestion API to make sure that it’s always returning the data you need, or you leverage a good API. You’re absolutely right. So, that is the code. That’s it. It’s really not that hard. It’s a couple of Graph calls plus a little bit of orchestration of, “I’m full crawling, I’m incremental crawling,” and these kinds of things. But it’s all built into the template that we give you. Awesome. Now, does it work? I have data, that’s great. What can I do with that data? Well, let’s go see. So, let me go here and let me create a new agent. And here, I’m going to create a new agent. I’m going to go straight to configure. Okay, I’m going to call it GitHub. And for anybody, while Seb does this if you need to know how to describe all that part, make sure you watch the other videos in the series because we already covered all of that. So, this time we’re only going to focus on what’s new right now. So, we’re going to go that it’s an AI that is expert in helping users with GitHub issues. And then the instruction is, “Help the user in a professional matter.”

“Identify important issues. When returning issues to the user, always use a table and the columns ID, title, and assigned. You have a little mistake there; instead of ID, you put ‘is.’ This one might not like it. If you want to be a little bit more specific, something that is really interesting is called a first shot. You can also give it a markdown example of what you would like your stuff to be. You can use the markdown syntax. It’s a little bit more on the dev side of things, but a lot of people are getting used to markdown these days, so markdown is a good way to represent that. In this case, does your letter casing matter? I see you put the ‘T’ in the capital. It should not be here; it’s just a regular dev. That’s how the property is really called in the schema, but it should not matter at all. That’s a good question. So here, we have knowledge, and this is where the knowledge will be coming from. Hey, look at that, I have my GitHub issues. We have it ready. So I have it ready, I click it, and then here I will not be adding anything else. I’ll just leave it like this. Now I can go here and I can say, ‘What are the issues related to authentication that I should be working on?’

Now it’s going to take the content from the ingested content in the graph connector, and it’s going to bring it back to you. Maybe it just needs a little bit of time since you just brought it in. No, it should not. So let me just create it and we’ll see how that plays out. Are there any limitations with the preview and this pane versus the real thing? It should not; both of them should be the same. I’m going to see how I could maybe help it a little bit once we’re here. So we’re going to restart here. So, ‘Get, list all authentication issues that are available in the GitHub agent repo.’ There we go. Okay, that worked well. So I wonder if ‘authentication’ didn’t get that exactly. There may be some reasons there, but now I am a little bit more specific. But it should not be too much of an issue. Now you see it rendered as the table that I asked it. Now I can go and just say, ‘Use these fields but use these titles to make it a little bit prettier,’ for instance. Now I have all of my bugs. So here, what is really great is that it already has the mobile app, so it’s already updated. If I click on that, now it’s a deep link directly inside my mobile app. And there we go.

Now, also look at the ‘API endpoint returning incorrect data.’ There’s no ‘auth’ as part of that title, but somewhere else in the description, it actually mentions that there’s an issue with auth. So that’s why I wanted to use this scenario; it’s because it brings things together. Here, I can also say something really interesting: ‘Which one should I prioritize?’ And this is where we believe generative AI is really interesting. Now it’s going to give you, like, ‘Given the two authentication issues, I recommend this one. This issue is critical as it directly impacts user access and security on the mobile app, which could lead to a poor user experience.’ So now it can do some summarization and help you along the way. So if we think about it, bringing the data in is just the easy part. Then afterwards, helping your agent will also make it better. Description instructions and so on will be better. Having some interesting starter prompts will also be interesting. In that case, I haven’t put any, but it’s going to be helpful for users to understand what they can do and how they should interact with the data. But in the end, we brought data from a third-party system, brought it into Copilot. Copilot was helpful to us, and now we’re ready to rock out of the box.

Does the agent know who you are? Can you ask it, ‘Are there any of them assigned to me?’ or ‘What are all my issues?’ or is that something extra we have to do? In this case, we can try it. It should actually look for my name. It looked for your name but not for my username. So what you could do is, as part of your ingestion, not only bring the username but bring the full profile of everybody. So in that case, it could actually match these things. But that’s something important that devs should look at because, especially for things like issues, people might wonder, ‘Okay, what’s assigned to me?’ Yes, exactly. But now, interesting, maybe I don’t have issues assigned to me in there. Well, you fixed all your issues. Exactly, thank you. You don’t have any bugs left, but that should absolutely work in that case. So that’s exactly how we think about it: bring the data in, have the user ask questions, and we really feel that there’s a lot of value there. That’s amazing, Seb. Thank you so much again for sharing all this amazing content. And again, this brings so many opportunities and ways that organizations can get more value out of Copilot and out of that investment they made. For everyone else, don’t forget to download the code; it will be in the description below. And I truly hope you enjoyed this video and found it valuable. You have our social media handles in the description below. Hopefully, you’re going to connect with us and make sure you check out the other episodes in the series. The playlist should pop up on your screen right about now. And of course, make sure you hit that subscribe button to get notified as soon as new content is available. See you in the next one. Cheers.”