SourceSync.ai
Research Date: 2026-02-23
Source URL: https://sourcesync.ai
GitHub Repository: https://github.com/scmdr/sourcesyncai-mcp (MCP server)
API Reference: https://sourcesync.ai/api-reference/authentication
Version at Research: SaaS platform (no version; REST API v1)
License: Proprietary SaaS (BYOC — Bring Your Own Cloud)
Overview
SourceSync.ai is a managed RAG (Retrieval-Augmented Generation) infrastructure platform that continuously syncs data from 15+ external sources into AI-ready knowledge bases. Teams configure ingestion once — connecting cloud storage, SaaS wikis, websites, or raw text — and SourceSync automatically cleans, chunks, and embeds content, keeping it fresh as sources change. It exposes the resulting knowledge base via a REST API with semantic and hybrid search, and through an MCP server for direct AI assistant integration.
Core Value Proposition: Eliminate custom RAG pipelines by providing auto-syncing, multi-source knowledge bases accessible over a standard API — with BYOC storage and embeddings so no data ever touches SourceSync's infrastructure.
Problem Addressed
| Problem |
Solution |
| Building RAG ingestion pipelines requires significant custom engineering |
Pre-built connectors for 15+ sources (cloud storage, SaaS, websites) eliminate bespoke ingestion code |
| AI knowledge becomes stale as source documents are updated |
Automatic sync keeps vector indexes fresh when upstream sources change |
| Semantic search alone misses keyword-important queries |
Hybrid search combines vector similarity with keyword matching, with configurable weights |
| Multi-tenant AI apps require isolated knowledge per customer |
Namespace + tenant ID model provides scoped data isolation without separate deployments |
| Storing AI data in vendor cloud creates compliance concerns |
BYOC model: file storage (S3-compatible) and vector DB (Pinecone) are customer-owned; SourceSync never retains data |
| Switching embedding models requires re-indexing everything |
Per-namespace embedding model config (OpenAI, Claude, Jina) allows different strategies per use case |
Key Statistics
| Metric |
Value |
Date Gathered |
| Pricing — Pilot |
$99/month (30M chars, 30k retrievals) |
2026-02-23 |
| Pricing — Pro |
$299/month (150M chars, 150k retrievals) |
2026-02-23 |
| Pricing — Team |
$999/month (750M chars, 750k calls) |
2026-02-23 |
| Enterprise |
Custom / unlimited |
2026-02-23 |
| Free tier |
14-day money-back on paid plans |
2026-02-23 |
| Connectors (live) |
Google Drive, Dropbox, OneDrive, Box, Notion |
2026-02-23 |
| Connectors (coming soon) |
Confluence, GitBook, S3, GCS, Azure Blob, Slack, Gmail, Salesforce, Zendesk, Intercom, SharePoint |
2026-02-23 |
| Scraping integrations |
Firecrawl, ScrapingBee, Jina |
2026-02-23 |
| Embedding providers |
OpenAI, Claude, Jina |
2026-02-23 |
| MCP Server |
sourcesyncai-mcp v1.0.11 on npm |
2026-02-23 |
| Notable customer |
SiteGPT.ai |
2026-02-23 |
Key Features
Namespace-Based Knowledge Isolation
- Each namespace is a fully independent knowledge base with its own storage, vector DB, and embedding model
- Namespace CRUD via REST API or MCP tools (
create_namespace, list_namespaces, etc.)
- Multi-tenant isolation via
X-Tenant-ID header — single API key can serve multiple end-users with separate data
Multi-Source Data Ingestion
- Direct ingestion: raw text, URL lists, sitemap crawl, full-website deep crawl (configurable depth/page limits)
- Cloud storage connectors (live): Google Drive, Dropbox, OneDrive, Box
- SaaS connectors (live): Notion
- Coming soon: Confluence, GitBook, S3, GCS, Azure Blob, Slack, Gmail, Outlook, Salesforce, ServiceNow, Zendesk, Intercom, SharePoint
- All ingestion is asynchronous —
get_ingest_job_run_status polls completion
- Connector reuse: system deduplicates OAuth connections across tenants automatically
Search
- Semantic search: vector similarity via
topK retrieval
- Hybrid search: configurable
semanticWeight + keywordWeight for precision/recall balance
- Retrieved documents include parsed text content URLs for downstream full-text access
Document Lifecycle Management
- Filter, update metadata, delete, and force-resync individual documents
fetchUrlContent retrieves raw parsed text from document content URLs
Bring Your Own Cloud (BYOC) Model
- Customer provides S3-compatible bucket, Pinecone index, and embedding API keys
- SourceSync processes and routes data but does not retain it
- Supports enterprise security and compliance requirements
MCP Integration
- Full MCP server (
sourcesyncai-mcp) maps all REST API operations to 28 MCP tools
- Drop-in for Claude Desktop, Cursor, Windsurf, Claude Code — zero SDK required
- See
research/mcp-ecosystem/sourcesyncai-mcp.md for detailed MCP tool inventory
Technical Architecture
Data Flow
┌──────────────────────────────────────────────────────────────────┐
│ SourceSync.ai Platform │
├──────────────────────────────────────────────────────────────────┤
│ │
│ Source Connectors Processing Pipeline │
│ ┌─────────────────┐ ┌───────────────────────────┐ │
│ │ Google Drive │ │ 1. Clean & extract text │ │
│ │ Dropbox / Box │──────► │ 2. Chunk content │ │
│ │ OneDrive │ │ 3. Embed (BYOE model) │ │
│ │ Notion │ │ 4. Store in BYOC storage │ │
│ │ URLs / Sitemaps │ └───────────┬───────────────┘ │
│ │ Raw Text │ │ │
│ └─────────────────┘ ▼ │
│ ┌─────────────────────┐ │
│ Customer Infrastructure: │ Namespace Index │ │
│ ┌──────────────────────┐ │ (Customer Pinecone) │ │
│ │ S3-compatible bucket │◄───│ │ │
│ │ Pinecone index │ └──────────┬──────────┘ │
│ │ Embedding model keys │ │ │
│ └──────────────────────┘ ▼ │
│ ┌─────────────────────┐ │
│ Consumer Layer: │ REST API / MCP │ │
│ ┌──────────────────────┐ │ - semantic_search │ │
│ │ Claude / Cursor / AI │◄───│ - hybrid_search │ │
│ │ Custom app via API │ │ - getDocuments │ │
│ └──────────────────────┘ └─────────────────────┘ │
└──────────────────────────────────────────────────────────────────┘
Namespace Configuration Stack
| Layer |
Customer-Provided |
| File Storage |
S3-compatible endpoint + credentials |
| Vector Store |
Pinecone API key, environment, index |
| Embedding Model |
OpenAI / Claude / Jina API key + model |
Authentication Model
- API Key: Bearer token in
Authorization header (per subscription)
- Tenant ID: Optional
X-Tenant-ID header for multi-tenant data isolation
- OAuth Connections: Per-connector OAuth flows for cloud storage / SaaS, with automatic connection reuse across tenants
Installation & Usage
REST API (No SDK)
# Semantic search against a namespace
curl -X POST https://api.sourcesync.ai/v1/search \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "X-Tenant-ID: tenant_XXX" \
-H "Content-Type: application/json" \
-d '{
"namespaceId": "namespace_XXX",
"query": "how do I configure embeddings?",
"topK": 5
}'
MCP Server (for AI assistants)
# Claude Desktop / Claude Code config
env SOURCESYNC_API_KEY=your_key npx -y sourcesyncai-mcp
{
"mcpServers": {
"sourcesyncai-mcp": {
"command": "npx",
"args": ["-y", "sourcesyncai-mcp"],
"env": {
"SOURCESYNC_API_KEY": "your_api_key",
"SOURCESYNC_NAMESPACE_ID": "your_namespace_id"
}
}
}
}
Example: Ingest a website, then search it
# 1. Ingest a website (async)
curl -X POST https://api.sourcesync.ai/v1/ingest \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{"namespaceId":"ns_XXX","ingestConfig":{"source":"WEBSITE","config":{"url":"https://docs.example.com","maxDepth":3}}}'
# 2. Poll for completion
curl https://api.sourcesync.ai/v1/ingest/job/run_XXX -H "Authorization: Bearer YOUR_API_KEY"
# 3. Hybrid search when complete
curl -X POST https://api.sourcesync.ai/v1/search/hybrid \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{"namespaceId":"ns_XXX","query":"getting started","hybridConfig":{"semanticWeight":0.7,"keywordWeight":0.3}}'
Relevance to Claude Code Development
Applications
- Skill reference grounding: Ingest all
SKILL.md files and references/ directories into a SourceSync namespace; query via MCP during skill execution for evidence-based answers
- Research directory search: Ingest
./research/**/*.md for semantic search across all research entries — surface relevant prior art when creating new entries
- Plugin documentation RAG: Index plugin README files and
plugin.json manifests for grounded plugin discovery and comparison
Patterns Worth Adopting
- BYOC / BYOE separation: The model of "we process, you own the data" cleanly separates vendor responsibility from data sovereignty — applicable pattern for plugin designs that handle user secrets
- Namespace isolation maps to plugin scope: SourceSync namespaces parallel Claude Code's plugin-scoped contexts; one namespace per project keeps retrieval signals clean
- Async ingestion with polling:
get_ingest_job_run_status pattern is reusable for any long-running knowledge base operation in agent workflows
Integration Opportunities
- Drop-in MCP context layer: Add
sourcesyncai-mcp to Claude Code's MCP server config to give all skills a live, searchable knowledge base without any RAG code
- Auto-synced research entries: Connect GitHub repository (via webhook or scheduled job) to SourceSync, so
./research/ changes are automatically indexed and searchable
- Competitive context grounding: Index competitor docs / changelogs in a namespace; query during skill creation to surface integration opportunities
Comparison with Related Tools
| Feature |
SourceSync.ai |
Microsoft GraphRAG |
Local Memory (local-memory) |
| Hosted/managed |
Yes (SaaS) |
No (self-hosted) |
No (local) |
| Auto-sync from sources |
Yes |
No |
No |
| Multi-source connectors |
Yes (15+) |
No (files only) |
No |
| Hybrid search |
Yes |
No (graph-based) |
No |
| BYOC storage |
Yes (required) |
No |
Yes |
| Knowledge graph |
No |
Yes |
No |
| Free tier |
No (paid only) |
Yes (open source) |
Yes (open source) |
| MCP server |
Yes |
No |
No |
| Multi-tenant |
Yes (X-Tenant-ID) |
No |
No |
References
Freshness Tracking
| Field |
Value |
| Last Verified |
2026-02-23 |
| Version at Verification |
REST API v1 (SaaS, versioned per release) |
| Next Review Recommended |
2026-05-23 |
Change Detection Indicators:
- Monitor connector roadmap for new live connectors (Confluence, GitBook, Slack are high-value)
- Track pricing changes — no free tier currently; watch for freemium tier announcement
- Verify vector store backend expansion beyond Pinecone
- Check for SDK releases (currently REST-only with no official SDK)
- Watch for self-hosted / open-core variant announcements
1---2name: problem-addressed-163description: SourceSync.ai is a managed RAG (Retrieval-Augmented Generation) infrastructure platform that continuously syncs data from 15+ external sources into AI-ready knowledge bases.4---5# SourceSync.ai67**Research Date**: 2026-02-238**Source URL**: <https://sourcesync.ai>9**GitHub Repository**: <https://github.com/scmdr/sourcesyncai-mcp> (MCP server)10**API Reference**: <https://sourcesync.ai/api-reference/authentication>11**Version at Research**: SaaS platform (no version; REST API v1)12**License**: Proprietary SaaS (BYOC — Bring Your Own Cloud)1314---1516## Overview1718SourceSync.ai is a managed RAG (Retrieval-Augmented Generation) infrastructure platform that continuously syncs data from 15+ external sources into AI-ready knowledge bases. Teams configure ingestion once — connecting cloud storage, SaaS wikis, websites, or raw text — and SourceSync automatically cleans, chunks, and embeds content, keeping it fresh as sources change. It exposes the resulting knowledge base via a REST API with semantic and hybrid search, and through an MCP server for direct AI assistant integration.1920**Core Value Proposition**: Eliminate custom RAG pipelines by providing auto-syncing, multi-source knowledge bases accessible over a standard API — with BYOC storage and embeddings so no data ever touches SourceSync's infrastructure.2122---2324## Problem Addressed2526| Problem | Solution |27|---------|----------|28| Building RAG ingestion pipelines requires significant custom engineering | Pre-built connectors for 15+ sources (cloud storage, SaaS, websites) eliminate bespoke ingestion code |29| AI knowledge becomes stale as source documents are updated | Automatic sync keeps vector indexes fresh when upstream sources change |30| Semantic search alone misses keyword-important queries | Hybrid search combines vector similarity with keyword matching, with configurable weights |31| Multi-tenant AI apps require isolated knowledge per customer | Namespace + tenant ID model provides scoped data isolation without separate deployments |32| Storing AI data in vendor cloud creates compliance concerns | BYOC model: file storage (S3-compatible) and vector DB (Pinecone) are customer-owned; SourceSync never retains data |33| Switching embedding models requires re-indexing everything | Per-namespace embedding model config (OpenAI, Claude, Jina) allows different strategies per use case |3435---3637## Key Statistics3839| Metric | Value | Date Gathered |40|--------|-------|---------------|41| Pricing — Pilot | $99/month (30M chars, 30k retrievals) | 2026-02-23 |42| Pricing — Pro | $299/month (150M chars, 150k retrievals) | 2026-02-23 |43| Pricing — Team | $999/month (750M chars, 750k calls) | 2026-02-23 |44| Enterprise | Custom / unlimited | 2026-02-23 |45| Free tier | 14-day money-back on paid plans | 2026-02-23 |46| Connectors (live) | Google Drive, Dropbox, OneDrive, Box, Notion | 2026-02-23 |47| Connectors (coming soon) | Confluence, GitBook, S3, GCS, Azure Blob, Slack, Gmail, Salesforce, Zendesk, Intercom, SharePoint | 2026-02-23 |48| Scraping integrations | Firecrawl, ScrapingBee, Jina | 2026-02-23 |49| Embedding providers | OpenAI, Claude, Jina | 2026-02-23 |50| MCP Server | `sourcesyncai-mcp` v1.0.11 on npm | 2026-02-23 |51| Notable customer | SiteGPT.ai | 2026-02-23 |5253---5455## Key Features5657### Namespace-Based Knowledge Isolation5859- Each **namespace** is a fully independent knowledge base with its own storage, vector DB, and embedding model60- Namespace CRUD via REST API or MCP tools (`create_namespace`, `list_namespaces`, etc.)61- Multi-tenant isolation via `X-Tenant-ID` header — single API key can serve multiple end-users with separate data6263### Multi-Source Data Ingestion6465- **Direct ingestion**: raw text, URL lists, sitemap crawl, full-website deep crawl (configurable depth/page limits)66- **Cloud storage connectors (live)**: Google Drive, Dropbox, OneDrive, Box67- **SaaS connectors (live)**: Notion68- **Coming soon**: Confluence, GitBook, S3, GCS, Azure Blob, Slack, Gmail, Outlook, Salesforce, ServiceNow, Zendesk, Intercom, SharePoint69- All ingestion is asynchronous — `get_ingest_job_run_status` polls completion70- Connector reuse: system deduplicates OAuth connections across tenants automatically7172### Search7374- **Semantic search**: vector similarity via `topK` retrieval75- **Hybrid search**: configurable `semanticWeight` + `keywordWeight` for precision/recall balance76- Retrieved documents include parsed text content URLs for downstream full-text access7778### Document Lifecycle Management7980- Filter, update metadata, delete, and force-resync individual documents81- `fetchUrlContent` retrieves raw parsed text from document content URLs8283### Bring Your Own Cloud (BYOC) Model8485- Customer provides S3-compatible bucket, Pinecone index, and embedding API keys86- SourceSync processes and routes data but does not retain it87- Supports enterprise security and compliance requirements8889### MCP Integration9091- Full MCP server (`sourcesyncai-mcp`) maps all REST API operations to 28 MCP tools92- Drop-in for Claude Desktop, Cursor, Windsurf, Claude Code — zero SDK required93- See `research/mcp-ecosystem/sourcesyncai-mcp.md` for detailed MCP tool inventory9495---9697## Technical Architecture9899### Data Flow100101```text102┌──────────────────────────────────────────────────────────────────┐103│ SourceSync.ai Platform │104├──────────────────────────────────────────────────────────────────┤105│ │106│ Source Connectors Processing Pipeline │107│ ┌─────────────────┐ ┌───────────────────────────┐ │108│ │ Google Drive │ │ 1. Clean & extract text │ │109│ │ Dropbox / Box │──────► │ 2. Chunk content │ │110│ │ OneDrive │ │ 3. Embed (BYOE model) │ │111│ │ Notion │ │ 4. Store in BYOC storage │ │112│ │ URLs / Sitemaps │ └───────────┬───────────────┘ │113│ │ Raw Text │ │ │114│ └─────────────────┘ ▼ │115│ ┌─────────────────────┐ │116│ Customer Infrastructure: │ Namespace Index │ │117│ ┌──────────────────────┐ │ (Customer Pinecone) │ │118│ │ S3-compatible bucket │◄───│ │ │119│ │ Pinecone index │ └──────────┬──────────┘ │120│ │ Embedding model keys │ │ │121│ └──────────────────────┘ ▼ │122│ ┌─────────────────────┐ │123│ Consumer Layer: │ REST API / MCP │ │124│ ┌──────────────────────┐ │ - semantic_search │ │125│ │ Claude / Cursor / AI │◄───│ - hybrid_search │ │126│ │ Custom app via API │ │ - getDocuments │ │127│ └──────────────────────┘ └─────────────────────┘ │128└──────────────────────────────────────────────────────────────────┘129```130131### Namespace Configuration Stack132133| Layer | Customer-Provided |134|-------|------------------|135| File Storage | S3-compatible endpoint + credentials |136| Vector Store | Pinecone API key, environment, index |137| Embedding Model | OpenAI / Claude / Jina API key + model |138139### Authentication Model140141- **API Key**: Bearer token in `Authorization` header (per subscription)142- **Tenant ID**: Optional `X-Tenant-ID` header for multi-tenant data isolation143- **OAuth Connections**: Per-connector OAuth flows for cloud storage / SaaS, with automatic connection reuse across tenants144145---146147## Installation & Usage148149### REST API (No SDK)150151```bash152# Semantic search against a namespace153curl -X POST https://api.sourcesync.ai/v1/search \154 -H "Authorization: Bearer YOUR_API_KEY" \155 -H "X-Tenant-ID: tenant_XXX" \156 -H "Content-Type: application/json" \157 -d '{158 "namespaceId": "namespace_XXX",159 "query": "how do I configure embeddings?",160 "topK": 5161 }'162```163164### MCP Server (for AI assistants)165166```bash167# Claude Desktop / Claude Code config168env SOURCESYNC_API_KEY=your_key npx -y sourcesyncai-mcp169```170171```json172{173 "mcpServers": {174 "sourcesyncai-mcp": {175 "command": "npx",176 "args": ["-y", "sourcesyncai-mcp"],177 "env": {178 "SOURCESYNC_API_KEY": "your_api_key",179 "SOURCESYNC_NAMESPACE_ID": "your_namespace_id"180 }181 }182 }183}184```185186### Example: Ingest a website, then search it187188```bash189# 1. Ingest a website (async)190curl -X POST https://api.sourcesync.ai/v1/ingest \191 -H "Authorization: Bearer YOUR_API_KEY" \192 -d '{"namespaceId":"ns_XXX","ingestConfig":{"source":"WEBSITE","config":{"url":"https://docs.example.com","maxDepth":3}}}'193194# 2. Poll for completion195curl https://api.sourcesync.ai/v1/ingest/job/run_XXX -H "Authorization: Bearer YOUR_API_KEY"196197# 3. Hybrid search when complete198curl -X POST https://api.sourcesync.ai/v1/search/hybrid \199 -H "Authorization: Bearer YOUR_API_KEY" \200 -d '{"namespaceId":"ns_XXX","query":"getting started","hybridConfig":{"semanticWeight":0.7,"keywordWeight":0.3}}'201```202203---204205## Relevance to Claude Code Development206207### Applications208209- **Skill reference grounding**: Ingest all `SKILL.md` files and `references/` directories into a SourceSync namespace; query via MCP during skill execution for evidence-based answers210- **Research directory search**: Ingest `./research/**/*.md` for semantic search across all research entries — surface relevant prior art when creating new entries211- **Plugin documentation RAG**: Index plugin README files and `plugin.json` manifests for grounded plugin discovery and comparison212213### Patterns Worth Adopting214215- **BYOC / BYOE separation**: The model of "we process, you own the data" cleanly separates vendor responsibility from data sovereignty — applicable pattern for plugin designs that handle user secrets216- **Namespace isolation maps to plugin scope**: SourceSync namespaces parallel Claude Code's plugin-scoped contexts; one namespace per project keeps retrieval signals clean217- **Async ingestion with polling**: `get_ingest_job_run_status` pattern is reusable for any long-running knowledge base operation in agent workflows218219### Integration Opportunities220221- **Drop-in MCP context layer**: Add `sourcesyncai-mcp` to Claude Code's MCP server config to give all skills a live, searchable knowledge base without any RAG code222- **Auto-synced research entries**: Connect GitHub repository (via webhook or scheduled job) to SourceSync, so `./research/` changes are automatically indexed and searchable223- **Competitive context grounding**: Index competitor docs / changelogs in a namespace; query during skill creation to surface integration opportunities224225### Comparison with Related Tools226227| Feature | SourceSync.ai | Microsoft GraphRAG | Local Memory (`local-memory`) |228|---------|--------------|-------------------|-------------------------------|229| Hosted/managed | Yes (SaaS) | No (self-hosted) | No (local) |230| Auto-sync from sources | Yes | No | No |231| Multi-source connectors | Yes (15+) | No (files only) | No |232| Hybrid search | Yes | No (graph-based) | No |233| BYOC storage | Yes (required) | No | Yes |234| Knowledge graph | No | Yes | No |235| Free tier | No (paid only) | Yes (open source) | Yes (open source) |236| MCP server | Yes | No | No |237| Multi-tenant | Yes (X-Tenant-ID) | No | No |238239---240241## References242243- [SourceSync.ai Website](https://sourcesync.ai) (accessed 2026-02-23)244- [SourceSync.ai Pricing](https://sourcesync.ai/pricing) (accessed 2026-02-23)245- [SourceSync.ai Connectors](https://sourcesync.ai/connectors) (accessed 2026-02-23)246- [API Reference — Authentication](https://sourcesync.ai/api-reference/authentication) (accessed 2026-02-23)247- [MCP Server Repository](https://github.com/scmdr/sourcesyncai-mcp) (accessed 2026-02-23)248- [MCP Server on Smithery](https://smithery.ai/server/@pbteja1998/sourcesyncai-mcp) (accessed 2026-02-23)249250---251252## Freshness Tracking253254| Field | Value |255|-------|-------|256| Last Verified | 2026-02-23 |257| Version at Verification | REST API v1 (SaaS, versioned per release) |258| Next Review Recommended | 2026-05-23 |259260**Change Detection Indicators**:261262- Monitor connector roadmap for new live connectors (Confluence, GitBook, Slack are high-value)263- Track pricing changes — no free tier currently; watch for freemium tier announcement264- Verify vector store backend expansion beyond Pinecone265- Check for SDK releases (currently REST-only with no official SDK)266- Watch for self-hosted / open-core variant announcements