AI-Powered Metadata in SharePoint: Structured and Free Form Extraction

Tired of begging users to add metadata to documents in SharePoint? What if AI could do the heavy lifting instead?

In this episode of the Ultimate SharePoint Content AI Deep Dive Series, Vlad and Microsoft MVP Gokan Ozcifci explore how SharePoint Content AI can extract metadata using structured and free form extraction models. Learn how to train AI Builder to recognize content from invoices, contracts, and more — all without forcing users to manually tag anything.

👀 We cover:
✔️ The difference between structured and free form models
✔️ When to use each method
✔️ Real-world example building a model with Uber invoices
✔️ How good is “good enough” for accuracy?
✔️ What to do when your model needs improvement

🎯 Whether you’re an information architect or just want to make SharePoint easier to use, this one’s for you.

📌 This is part of our Ultimate SharePoint Content AI series — check out the full playlist to go even deeper.

Watch my 100+ courses on Pluralsight

Video Summary

  • Metadata matters—but users still skip it.
    Despite years of reminders, most users don’t add metadata in SharePoint. Now, AI can automate that process and make everyone’s life easier.
  • Choose the right AI model: structured vs. freeform.
    Use structured models for predictable formats like invoices, and freeform models for less structured content like contracts. Both rely on AI Builder and Power Platform.
  • Creating a model is simple and user-friendly.
    Upload a few sample files, highlight the data you want to extract, and let AI Builder do the rest. Just remember: you need at least five documents to train the model.
  • Accuracy depends on consistency.
    A model with a 60% confidence score still worked well, because the documents were consistent. For best results, aim for 95–100% accuracy and use similar layouts.
  • You can always improve and reuse your model.
    Not satisfied? Retrain your model in Power Apps or Power Automate. You can also apply the same model across multiple SharePoint libraries.

For more information, read the transcript blog below, or watch the video above!

Transcript

One of the biggest pain points that SharePoint users have had for the past two decades is adding metadata to the documents they upload to SharePoint. Of course, information managers have told them time and time again why it was important and how it could benefit their usage in SharePoint, but they still didn’t do it.

What if I told you there’s a way to make both users and information managers happy by letting AI do the hard work? In this video, I’m joined by the amazing Microsoft MVP Gokan Ozcifci, who will deep dive into freeform and structured extraction for SharePoint documents.

“Gokan, welcome back!”
“Hey, thank you, Vlad. I just love your introductions—the way you explain things. I’m pretty sure all the information managers listening to us will say, ‘Oh my gosh, Vlad, you’re so right.’ Everyone knows it’s important, but no one takes care of metadata. Today’s video is all about how we can fix this.”

“My name is Gokan Ozcifci, I’m from Brussels, Belgium, and I’m extremely happy to be here to talk about SharePoint content in your videos, Vlad.”

This is the fourth episode in our SharePoint Content AI series. For those of you who have been following the series so far, welcome back! For those of you watching for the first time—welcome! We really hope you enjoy the video, and after you’re done with this one, make sure to check out the previous episodes to learn more about SharePoint Content AI.

Now let’s get started.
“Gokan, I’ve heard terms like structured and unstructured, but I have no idea what they mean. Can you go through that?”
“Yeah, for sure. We have a bunch of naming conventions for the AI models in SharePoint Content AI, like semi-structured, structured, and unstructured. Unstructured is the SharePoint way of doing things, and we’ll cover that in another video. In this video, we’ll focus our energy on the structured AI models, meaning the content has some structure—could be semi-structured or fully structured—but there’s enough structure to extract data from it. Tables, invoices, contracts—these are examples of structured content where you can easily pinpoint and extract data.”

“Unstructured content is where the data could be anywhere—like a company name in the upper-left corner, or on the 16th page. Freeform or semi-structured means we kind of know where the data is—on this side or that side of the document.”

“If I have to deep dive into structured models, we use structured document processing for structured file formats like forms or invoices. The biggest thing is that it relies on the AI Builder methodology, meaning we’re consuming knowledge from the Power Platform. That also means we need a Power Platform subscription. We’ll see that in a couple of seconds, but that’s one of the biggest drawbacks—having to pay for it.”

“Then we have freeform models, which are for formats with no set structure—like letters, contracts, or statements of work. For example, a 15-page contract should go into the freeform model. It also relies on AI Builder. From an end-user perspective, it’s the same—same options, same screens. The only difference is deciding whether to go with freeform or structured, depending on which is better for your content.”

“When we look at the screen, you’ll see a table comparing unstructured and structured models. One important thing for structured document processing is the transactional cost. Unstructured uses a pay-as-you-go model but structured requires AI Builder credits. There’s a yearly commitment for those credits. If you have Power Automate Premium or Power Apps Premium, you might already have AI Builder credits included—so for some of you, it might be free.”

“The main advantage is that it’s extremely easy to use—and I’ll show that in a couple of seconds. Oh, and by the way, I’ve run out of slides—and you know I don’t really like slides—so let’s go into the demo.”

“I have a SharePoint site here, and on the left side, a document library. When I go into the document library and click the three dots, I see ‘Classify and Extract.’ For some reason, it refreshed—it’s okay, SharePoint wanted to make sure it’s fresh for you. Under ‘Classify and Extract,’ you’ll see all the available AI models for the document library. You can click ‘Apply’ or see if any model has already been applied. We have none, so we’ll create a new one.”

“As I said, you have the freeform and structured options. From an end-user perspective, there’s no difference. Let’s go with freeform. It gives you a few examples of what the model can do and the supported file types—PDFs and images only. No Word documents or PowerPoint—just PDFs and images.”

“Click ‘Next,’ give it a name—let’s say ‘Gokan’s Travel Expenses.’ Under advanced settings, you can choose a content type—either create a new one or select an existing one. From an information architect’s perspective, using an existing one is best. People who invest in content types love them. You can also select a sensitivity label and retention policy. I have one called ‘Neoxy – Not to be deleted.’ If you select that, your data can’t be deleted. For demo purposes, I’ll select ‘None’ and click ‘Create.’”

“Now we leave the SharePoint experience and enter the AI Builder experience. It’s completely different. Here, you define the information you want to extract—text field, number field, date field, checkbox, or table. I’ve never used a table before—just being honest. I usually use text and number fields.”

“Let’s go with a text field. It asks what you want to extract—let’s say ‘Address.’ Then I’ll create another one—‘Destination.’ You can add as many as you like. These will create columns in your document library and populate them with the extracted data.”

“Click ‘Next.’ Now it asks if you want to create a collection to add your invoices. I’ll create one called ‘Uber.’ A collection is like a container for your documents. If you have invoices from Uber, Lyft, or taxis, you can create separate collections. This helps the AI understand and apply the model.”

“Click ‘Uber,’ then ‘+’ to add documents. You can upload up to 1 GB or 50,000 pages. This is great for large-scale processing. There’s no strict minimum, but I always recommend at least five documents. Let’s try uploading one—oh, it says the minimum is five. I didn’t know that because I always upload five!”

“Upload your five documents, and you’ll see them in the collection. Click ‘Next.’ Now you get previews of each document. Highlight the fields you want to extract—‘Address’ and ‘Destination.’ Do this for each document. Since they’re structured, the data is usually in the same place.”

“Once done, click ‘Next.’ You’ll get a summary—owner, extraction type, number of collections, number of documents, and metadata fields. Click ‘Train.’ Training can take 20 minutes to a few hours. Since we’re filming this at the Microsoft 365 Community Conference in Las Vegas, let’s go attend a few sessions while it trains.”

“We’re back! The background looks different—we’re now in Vienna. The model took 15–20 minutes to train, but we didn’t get a chance to finish the demo until now. Let’s test it.”

“The model has a 60% confidence score—not great. The address field scored poorly, and the value field only got 80%. It’s not always perfect. You might need to tweak the model—add more files, refine your selections, and retrain.”

“To modify the model, go to Power Apps or Power Automate, then the AI Hub. You’ll see your trained model there. Click to edit, add more files, or adjust the fields. Aim for 95–100% confidence.”

“You can also test the model in SharePoint and apply it to multiple libraries. I uploaded more invoices, and within seconds, the AI extracted the values. It did a phenomenal job because the files were consistent. If your files vary in layout or text, the model might fail.”

That’s it for the AI Builder structured document processing model. In the next videos, we’ll explore prebuilt models and autofill—easier ways to automate metadata extraction in SharePoint. Thanks for watching, and don’t forget to like, subscribe, and check out the next video in the series!