
Self-Maintaining Intune App Catalogue with AI
How I built a self-maintaining Intune app catalogue: Graph API pulls inventory, AI classifies it, and surfaces your MSIX migration shortlist automatically.
In Part 4 of the Windows Endpoint Management series, I described the catalogue problem in about 200 words and waved my hands at a solution. Graph API, Power Automate, AI classification: the building blocks were named but not assembled. This post is the 4,000-word version where we actually build the thing.
Someone asks "what applications do we have deployed?" and the answer takes three weeks of archaeology through Intune portals, spreadsheets of varying vintage, and conversations with people who left the organisation two years ago. The catalogue that was supposed to answer this question died quietly in a SharePoint list that nobody updated after March.
This post fixes that. Spreadsheet discipline has already been tried, and it didn't hold. A pipeline that maintains the catalogue automatically, enriches it with AI, and surfaces the outputs you need for migration and governance decisions does the job instead.
Why manual catalogues die
Part 3 introduced the application catalogue as "a centralised, authoritative record of every application in your estate." The fields were clear: application name, version, packaging format, owner, deployment method, ring assignments, dependencies, and lifecycle status. The advice was sound: export your Intune app list (see Graph API for Architects, which covers the endpoints in detail), enrich it, and maintain it.
Here's what actually happens.
Someone exports the list and spends an afternoon adding columns for owner and lifecycle status, then presents it at a team meeting, where everyone agrees it's valuable. For about three weeks, new applications get added and versions get updated. Then a deadline hits, the catalogue update gets skipped once, then twice, and within a quarter it's a historical artefact, accurate as of the date it was last touched, misleading for everything after.
Blaming laziness misses the real cause: manual catalogue maintenance takes ongoing effort with no immediate reward. The person updating the spreadsheet doesn't benefit from doing it. The benefit accrues later, to a different person, when they need to answer a question the catalogue can answer. This asymmetry between effort and benefit is why every manual catalogue eventually fails.
The fix is obvious in retrospect: stop asking humans to do what machines can do. Intune already holds this data, what's deployed and what's installed. It just needs extracting, correlating, and enriching, and every one of those steps can be automated.
Architecture
DATA SOURCES
PROCESSING
ACTIONABLE OUTPUTS
The pipeline has three tiers.
The pipeline collects data first: the Intune Graph API exposes two datasets, managed apps (what you intentionally deploy) and detected apps (what's actually installed). A scheduled process pulls both.
The raw data then goes through processing: normalised, deduplicated, and sent to an AI model for classification. The AI correlates the two datasets, identifies gaps, and enriches each entry with metadata that flags whether the app is orphaned, an MSIX candidate, or past end of life.
The output lands in a SharePoint list that is the living catalogue. Views and filters surface the actionable items, orphaned software to investigate, MSIX conversion candidates to prioritise, and applications flagged for retirement.
Power Automate fits because the audience for this pipeline is IT operations teams, and most of them already have Power Automate as part of their Microsoft 365 licensing. The HTTP connector with Entra ID authentication needs a Power Automate Premium licence; beyond that, there's no Azure subscription, no custom infrastructure, and no deployment pipeline to maintain. If you prefer a configuration-as-code approach, the companion repository also provides a pure PowerShell implementation you can run from a local machine, Task Scheduler, Azure Automation, or a CI/CD pipeline.
What Graph API gives you
Two endpoints do most of the work.
Managed apps
GET /deviceAppManagement/mobileApps returns every application that your organisation has configured in Intune: Win32 packages, MSIX bundles, Store apps, WinGet deployments. The @odata.type property tells you the packaging format, which matters for the MSIX migration analysis. The endpoint also exists on v1.0, but Microsoft documents the winGetApp type on beta only, so the sample below stays on beta to identify WinGet apps.
# Fetch managed apps with packaging format
$uri = "https://graph.microsoft.com/beta/deviceAppManagement/mobileApps?`$top=100"
$response = Invoke-RestMethod -Uri $uri -Headers $headers -Method Get
# The @odata.type reveals the format
# #microsoft.graph.win32LobApp → Win32
# #microsoft.graph.windowsAppX → MSIX
# #microsoft.graph.winGetApp → WinGet
# #microsoft.graph.microsoftStoreForBusinessApp → Store (legacy; retired)Detected apps
GET /deviceManagement/detectedApps returns what's actually installed across your estate, regardless of how it got there: the ground truth, not the intended state. It lists every application the Intune management extension has discovered on managed devices, complete with a device count showing how widespread each installation is.
# Fetch detected apps with device counts
$uri = "https://graph.microsoft.com/v1.0/deviceManagement/detectedApps?`$top=100"
$response = Invoke-RestMethod -Uri $uri -Headers $headers -Method Get
# Each entry includes:
# displayName, version, publisher, platform, deviceCountThe gap
The managed list tells you what should be there. The detected list tells you what is there. The gap between them is your entire problem space:
| Scenario | Managed | Detected | Classification |
|---|---|---|---|
| App deployed and found on devices | Yes | Yes | Managed (healthy) |
| App in Intune but not detected anywhere | Yes | No | Check assignments |
| App found on devices but not in Intune | No | Yes | Orphaned (investigate) |
| Win32 app with no complex install logic | Yes | Yes | MSIX candidate |
That third row (installed but not managed) is where the security and compliance risk lives, and where shadow IT shows up: users installing software outside Intune, legacy deployment tools still pushing packages, vendor bloatware that came with hardware. Detected apps surfaces all three, which is what makes shadow IT visible in the first place.
Both endpoints paginate via @odata.nextLink. For estates with more than a few hundred apps, you'll need to follow the pagination chain. The repo scripts handle this automatically.
Building the sync flow
The flow runs daily (or on demand) and executes six stages:
- Scheduled trigger. A daily recurrence at 06:00 UTC. For most organisations, daily is enough, since application inventory doesn't change hour by hour. If you need tighter feedback, you can supplement the schedule with event-driven triggers when new apps are deployed.
- Fetch managed apps. An HTTP action calls the managed apps endpoint, authenticating with a Microsoft Entra ID app registration that holds the
DeviceManagementApps.Read.Allpermission (read-only; the pipeline never modifies your Intune tenant), and follows@odata.nextLinkuntil it has every page. - Fetch detected apps. A second HTTP action runs in parallel, calling the detected apps endpoint with the
DeviceManagementManagedDevices.Read.Allpermission and following its own@odata.nextLinkchain. - Normalise. Raw Graph API data needs cleaning before classification. Application names contain version numbers, for example "Google Chrome 121.0.6167.85" versus "Google Chrome". Publishers use inconsistent casing, "Microsoft Corporation" versus "microsoft corporation". The same application can also appear multiple times across different detection contexts. This step standardises names, deduplicates entries, and merges the managed and detected datasets into a single view.
- AI classify. The normalised dataset goes to an AI model with a carefully designed system prompt. Any model that supports structured JSON output works: OpenAI, Microsoft Foundry (formerly Azure OpenAI), Claude, Gemini, or a local model via Ollama. The AI understands context beyond the name: that "Microsoft Visual C++ 2019 Redistributable" is a runtime dependency, that "Notepad++" with 187 devices and no Intune deployment is orphaned, and that "Google Chrome" as a Win32 package is a strong MSIX conversion candidate. More on the classification logic in the next section.
- Write catalogue. Results sync to a SharePoint list using delta logic: match on a composite key (application name and publisher), update existing entries, create new ones, and flag entries that have disappeared since the last run. The pipeline never deletes; a disappeared app gets its lifecycle status updated to "Under Review" for a human decision. The Power Automate flow definition in the repo creates SharePoint items without matching existing ones, so it has no delta sync; Update-Catalogue.ps1, the PowerShell script in the repo, implements the logic above.
The AI classification step
The AI classification step is the heart of the pipeline, the part that transforms raw inventory data into actionable intelligence.
Five categories
Each detected application receives exactly one classification.
Managed means the detected app matches a managed Intune deployment. This is the healthy state: you're intentionally deploying this software and it's landing on devices as expected, so no action is required.
An orphaned app is installed somewhere in the estate but isn't deployed or managed through Intune. It arrived through legacy deployment tools, user self-install, vendor pre-installation, or some other uncontrolled channel. Every orphaned app raises a compliance and security question that needs answering.
Ownership goes missing in a different way: the app is a managed deployment, but nobody is accountable for patching it, reviewing its business justification, or deciding when to retire it. That governance gap compounds over time.
An MSIX candidate is currently packaged as Win32 and looks suitable for conversion: no kernel-mode drivers, no global registry writes outside the package, no dependencies on install-time actions that mutate system state. These are the easiest conversion targets, the ones Part 3 argued you should tackle first. The AI also assigns an MSIX readiness score from 1 (unlikely) to 5 (trivial).
An app heading for retirement looks end-of-life, is superseded by a newer product, or has a device count near zero, suggesting nobody needs it any more. These are candidates for a governed retirement process.
The prompt matters
The quality of classification depends entirely on the system prompt. A naive "classify these apps" instruction produces mediocre results. The system prompt in the repo encodes domain knowledge that the AI wouldn't otherwise have:
- Framework runtimes (.NET, Visual C++ Redistributable) are dependencies, not standalone applications
- An app detected on more devices than its managed deployment count suggests shadow installations
- Version comparison with vendor lifecycle data identifies end-of-life software
- MSIX readiness assessment requires understanding which installer characteristics are incompatible with containerised execution
Structured output is essential. The prompt requests JSON matching a defined schema, not prose. Every classification includes the category, a confidence reason, the matched managed app (if any), and for MSIX candidates, the readiness score. Machine-readable output means the pipeline can act on results directly (creating SharePoint items, generating reports, populating migration backlogs) without human parsing.
Two implementation paths
Power Automate plus AI uses the HTTP connector to call your AI provider's API directly from the flow. The system prompt, few-shot examples, and app data go in the request body, and structured JSON comes back. For organisations already in the Power Platform, this keeps everything in one place.
PowerShell plus API uses the Invoke-AppClassification.ps1 script in the repo, which handles batching (splitting large estates into chunks that fit within token limits), retry logic, and structured output parsing. This path gives you more control and fits naturally into CI/CD pipelines or Azure Automation.
Both paths use the same prompts and produce the same output schema. The repo defaults to OpenAI's chat completions format, but most providers either use the same format natively (Microsoft Foundry, Ollama) or offer compatible endpoints (Gemini). For providers with different APIs, like Anthropic's Claude, you'll need to adjust the HTTP call, but the prompts and schema are provider-agnostic. Choose based on what your organisation already has access to.
The catalogue as a living asset
The SharePoint list is the single source of truth. Its schema has two kinds of fields:
Automated fields, populated and updated by the pipeline on every run:
| Field | Source | Description |
|---|---|---|
| Application Name | Graph API | What the app is called |
| Publisher | Graph API | Who makes it |
| Version | Graph API | What's currently installed |
| Classification | AI | Managed, orphaned, unowned, MSIX candidate, or retirement |
| Device Count | Graph API | How many devices have it |
| Packaging Format | Graph API | Win32, MSIX, Store, etc. |
| MSIX Readiness | AI | Conversion difficulty score (1 to 5) |
| Classification Reason | AI | Why the AI classified it this way |
| Last Sync Date | Pipeline | When this entry was last refreshed |
Manual fields, needing human judgement:
| Field | Description |
|---|---|
| Owner | Who's responsible for this application |
| Business Justification | Why the organisation needs it |
| Lifecycle Status | Active, Under Review, Deprecated, Retired |
| Notes | Context that doesn't fit elsewhere |
The pipeline never overwrites manual fields. When it updates an entry, it refreshes the automated columns and leaves human-entered data intact. That split keeps the catalogue maintainable: the pipeline automates the tedious part, inventory and classification, and people supply what only they can, ownership decisions, business context, and governance judgement.
Delta sync means the pipeline matches incoming data against existing entries using a composite key (application name and publisher). New apps become new rows. Known apps get updated columns. Apps that disappear from the detected inventory have their lifecycle status flagged as "Under Review" rather than being deleted, because "no longer detected" doesn't necessarily mean "should be removed from the catalogue."
From catalogue to migration backlog
This is the payoff. The catalogue generates a queue of work: MSIX candidates to convert, orphaned apps to investigate, and retirement candidates to clean up.
Filter the catalogue view to "MSIX Candidates", sort by readiness score (descending) and device count (descending), and you have a prioritised migration backlog. The apps at the top (high readiness, high device count) deliver the most value for the least effort. These are your first batch.
Part 3 argued that MSIX should be the default packaging format, with Win32 reserved for genuine exceptions. The usual objection was: "We don't know which of our Win32 apps can actually be converted." This pipeline answers that question with a specific list, complete with readiness scores, backed by AI analysis of each application's characteristics, rather than a guess at "maybe most of them."
A typical enterprise estate with 200 or more applications usually surfaces 40 to 80 MSIX candidates, of which 15 to 30 score 4 or 5 on the readiness scale, meaning straightforward or trivial conversions. That's a concrete starting point: a specific backlog, with scores attached, that a packaging team can work through instead of a vague plan to look into MSIX someday.
The other outputs are equally actionable. Orphaned apps need a security review: who installed this, why, and whether it should be managed or removed. Unowned apps need owners assigned, justification documented, or a flag for retirement. Retirement candidates go into a cleanup project that removes the deployment, reclaims licences, and reduces attack surface.
Each of these outputs can be pushed downstream into Planner tasks, Azure DevOps work items, ServiceNow tickets, or whatever work management system your organisation uses. The pipeline both flags the problem and creates the ticket to fix it.
Running it in practice
First run
Expect noise. The first run against a real estate will over-classify, under-classify, and occasionally hallucinate; that's normal. Common issues:
- The AI classifies framework runtimes as orphaned, because they aren't in the managed app list as standalone deployments. The few-shot examples in the repo handle the most common ones, but you'll need to add examples specific to your estate.
- Name mismatches happen where the detected name differs from the managed app name, for example "Adobe Acrobat Reader DC" versus "Adobe Acrobat Reader" versus "AcroRd32".
- The AI overestimates MSIX readiness for apps that look simple from their name but have complex installation requirements.
Review the first run manually. Correct misclassifications by updating the few-shot examples in prompts/few-shot-examples.json. Each correction improves subsequent runs. After two or three iterations, classification accuracy stabilises.
Steady state
Once tuned, the pipeline runs daily with minimal attention. What you monitor:
- A weekly digest: new orphaned apps (someone installed something new), classification changes (an app's risk profile shifted), and apps that disappeared from the estate.
- Catalogue coverage: the percentage of detected apps that have been classified. This should sit at or near 100%.
- Owner coverage: the percentage of managed apps that have an assigned owner, your governance maturity metric. The pipeline highlights the gap, but closing it needs human action.
- Migration progress: how many MSIX candidates remain. Track this over time as your packaging team works through the backlog.
Where this goes next
This post built the data foundation. The catalogue is a living, self-maintaining record of your application estate, and it now answers questions the SharePoint list never could.
The next post in this series will take two of the catalogue's outputs, orphaned and unmanaged applications, and build an automated compliance remediation workflow. When an orphaned app appears, the pipeline will determine the risk, draft a remediation plan, and route it to the right team for action, instead of only raising a flag.
If you've read Part 4 of the main series, you'll recognise this as one step toward the agentic loop of monitor, detect, plan, act, and verify. The catalogue is the detect layer: it tells you what you have. Compliance remediation is the plan and act layer: it does something about what you've found.
That's for next time. For now: clone the repo, point it at your tenant, and see what falls out. The first time you run the pipeline and it tells you that 23% of your installed software isn't managed by Intune, you'll understand why this catalogue matters and why it needs to maintain itself.
This is Part 1 of the Practical AI for Endpoint Management series: hands-on builds that turn the concepts from the Windows Endpoint Management series into working systems.
Get the next one by email
Long, specific write-ups on M365 architecture and security, worked out against real tenants rather than summarised from documentation. Sent rarely, and only when it is worth your time.