GigaOmResearch
The New GigaOm Radar is coming soon!Learn more
Commissioned byMicrosoft
TCO & Benchmark

Evaluating Microsoft 365 Copilot’s Work IQ vs ChatGPT with Microsoft 365 Connectors

Enterprises are rapidly adopting AI-powered tools to enhance workforce productivity, decision-making, and organizational agility. These systems are evolving beyond standalone large language models into connected reasoning engines.

Executive Summary

Enterprises are rapidly adopting AI-powered tools to enhance workforce productivity, decision-making, and organizational agility. These systems are evolving beyond standalone large language models into connected reasoning engines capable of retrieving, synthesizing, and acting on data distributed across enterprise environments, including emails, meetings, chats, and documents.

Within the Microsoft ecosystem, this capability is enabled by what Microsoft refers to as Work IQ—the intelligence layer that allows Microsoft 365 Copilot to deeply understand a user and their organization. Work IQ does this by understanding the data, relationships, context, and workflows that each form the basis of a user’s way of working.

Work IQ combines three core elements: secure access to organizational data across Microsoft 365 and connected business systems; contextual understanding of how work happens across people, projects, and communications; and the ability to apply skills and tools to take action on that information. As a result, Copilot can generate responses that are not only accurate, but contextually relevant and aligned to how work is actually performed within the organization.

As these tools become more deeply embedded in day-to-day operations, the stakes extend beyond productivity. The ability to securely access and reason over enterprise data—while adhering to security and governance tooling such as Microsoft Purview, sensitivity labels, and data loss prevention (DLP) policies—becomes a critical requirement for enterprise adoption.

The central question is whether integration architecture—native versus connector-based—produces meaningful differences in accuracy, completeness, and adherence to enterprise data governance when both platforms are pointed at the same organizational data.

To assess this, we conducted an evidence-based comparison between Microsoft 365 Copilot and OpenAI ChatGPT Business (connected to the Microsoft 365 environment via Connectors). The evaluation was performed in a controlled environment simulating a real-world Microsoft 365 deployment, including emails, Teams channels and chats, calendar events, transcribed meetings, SharePoint documents, and OneDrive files—some governed by Microsoft Purview sensitivity labels and DLP protections.

Across 25 workplace prompts and eight enterprise personas, the benchmark evaluates:

  • Accuracy in retrieving and synthesizing information from enterprise data sources
  • Completeness of responses across stakeholders, deliverables, and timeframes
  • Adherence to enterprise data governance policies
  • Relative strengths across key categories of workplace tasks, including personal, team, and company-wide workflows

Key Findings

The results indicate that, in enterprise use cases, integration architecture is more impactful for both efficacy and performance than underlying model capability. Microsoft 365 Copilot demonstrates a structural advantage in enterprise scenarios that require accessing and synthesizing information across multiple data sources. Copilot consistently delivered more complete, accurate, and context-aware responses, particularly in workflows involving Teams, Outlook, SharePoint, and OneDrive. Its native integration with the Microsoft 365 enterprise graph enabled stronger performance in organizational context, task identification, and cross-application reasoning, while also reducing hallucination risk in fact-based queries.

In contrast, ChatGPT performed competitively in isolated or content-transformation tasks where all required information was provided directly. However, its reliance on connectors introduced variability in data access, with missed sources and intermittent reliability issues impacting completeness and consistency. These results underscore that performance differences are driven less by core model capability and more by the depth of enterprise integration and the ability to reliably traverse organizational data.

Microsoft 365 Copilot scored 9/10 or higher in accuracy for 21 of 25 prompts, while ChatGPT accomplished this for 12 of 25. For comprehensiveness, Copilot scored 9/10 or higher for 17 of 25, while ChatGPT managed just eight.

Microsoft 365 Copilot passed the two prompts that were specifically designed to test for Trust, checking for compliance with Purview DLP policies, while ChatGPT failed both of these prompts.

Read the full TCO & Benchmark

Sign in to read the rest of this report, including the full analysis and the analyst's verdict.

The full report covers

  • Market Context and Background
  • Results Overview
  • Test Methodology
  • Benchmark Results
  • Conclusion
  • Appendix