Development Standards
This document introduces various standards and best practices in the project development process to help the team maintain code quality and collaboration efficiency.
🚀 TL;DR (Core Principles)
Quick Start for Newcomers (2 Steps)
make dev-setup # One-click dev environment setup (sync deps + install hooks)
Core Conventions
📦 Dependency Management
Use uv add/remove to manage dependencies, avoid direct pip install to maintain consistency of dependency lock files
🎨 Code Style
Pre-commit checks run automatically on commit (black/ruff/isort) to keep code style consistent
⚡️ Full Async Architecture
Single Event Loop, use async/await for I/O operations, discuss with development lead before using threads/processes
🚫 No I/O in Loops
Prohibit database access and API calls in for loops, use batch operations instead
🕐 Timezone Awareness
All time fields must carry timezone information. Input without timezone is treated as the timezone configured by environment variable TZ (default UTC). Do not use datetime.datetime.now(), must use utility functions from common_utils/datetime_utils.py
📥 Import Standards
- PYTHONPATH management: Project module import starting paths (src/tests/demo etc.) need unified management, communicate with development lead before changes
- Prefer absolute imports (e.g.
from core.memory import MemoryManager), avoid relative imports (e.g.from ...core import)
📝 init.py Standards
Not recommended to write any code in __init__.py, keep it empty
🌿 Branch Standardsdev for daily development, release/YYMMDD for versioned releases, long/xxx for long-term feature development, hotfix for emergency fixes
🔀 Unified Branch Merge Handling
Merging long/xxx to dev, cutting release from dev, merging release back to dev needs to be handled uniformly by development or operations lead
📤 MR Standards
- Keep code commits small, iterate quickly, avoid submitting too much code at once
- Each commit should be runnable, do not submit work-in-progress or broken code
- Data migration scripts, dependency changes, infrastructure code changes, merging release branches must go through Code Review
💾 Data Migration Standards
For new features involving data fixes or Schema migration, discuss feasibility and implementation timing with development and operations as early as possible
🏛️ Data Access Standards
All database, search engine and other external storage read/write operations must be converged to infra layer repository methods. Direct calls in business layer are prohibited
🎯 Minimal Changes
Minimize code changes when implementing requirements, avoid large-scale refactoring, prioritize incremental development. Do not over-engineer, keep it simple, efficient, and maintainable
💬 Comment Standards
Always add sufficient comments (function-level + step-level), ensure reviewers can quickly understand code intent
📖 API Documentation Sync
When modifying API interfaces, must synchronize updates to API documentation comments, schema definition files, and auto-generated documentation files
📄 Documentation Standards
Use markdown format, place in docs directory. Small issues don't need documentation, just add comments in code
🌍 Internationalization Standards
For international team communication and collaboration, code comments and documentation should be written in English
📖 Quick Navigation
- Don't know how to install dependencies? → Dependency Management Standards
- Need database/middleware configuration? → Development Environment Configuration Standards
- Always getting errors before commit? → Code Style Standards
- How to write code comments? → Comment Standards
- What to do after changing API? → API Specification Sync
- Not sure if you can use threads? → Async Programming Standards
- Can I do database queries in loops? → Prohibit I/O Operations in for Loops
- How to handle time fields? → Timezone Awareness Standards
- Where should database queries be written? → Data Access Standards
- Import path errors? → Import Standards
- How to name module introduction files? → Module Introduction File Naming
- Don't know which branch to use? → Branch Management Standards
- How to submit code/Need to submit MR? → MR Standards
- Need data migration? → Data Migration and Schema Change Process
📋 Table of Contents
- TL;DR (Core Principles)
- Dependency Management Standards
- Development Environment Configuration Standards
- Code Style Standards
- Comment Standards
- API Specification Sync
- Documentation Standards
- Async Programming Standards
- Timezone Awareness Standards
- Data Access Standards
- Import Standards
- Module Introduction File Naming
- Internationalization Standards
- Branch Management Standards
- MR Standards
- Code Review Process
📦 Dependency Management Standards
Using uv for Dependency Management
💡 Important Note: Recommended to use uv for dependency management
The project uses uv as the dependency management tool. It's recommended to avoid using pip install directly for the following reasons:
- Dependency versions may be inconsistent
uv.lockfile cannot be automatically updated- Team member environments may differ
- May affect production environment deployment
Correct Operations
1. Install/Sync Dependencies
# Sync all dependencies (first install or after updates)
uv sync --group dev-full
2. Add New Dependencies
# Add production dependency
uv add <package-name>
# Add development dependency
uv add --dev <package-name>
# Specify version
uv add <package-name>==<version>
3. Remove Dependencies
uv remove <package-name>
4. Update Dependencies
# Update all dependencies
uv sync --upgrade
# Update specific dependency
uv add <package-name> --upgrade
Related Documentation
For detailed dependency management guide, refer to: project_deps_manage.md
🔧 Development Environment Configuration Standards
Environment Configuration Description
The project depends on various databases and middleware. To ensure consistency and security of development environments, these configurations are uniformly managed and distributed by the operations team.
Configuration Items
Development environment typically needs the following configurations:
Database Configuration
- MongoDB connection information
- PostgreSQL connection information
- Redis connection information
Middleware Configuration
- Kafka connection configuration
- ElasticSearch connection configuration
- Other message queues or cache services
Third-party Service Configuration
- API keys and access credentials
- Object storage configuration
- Other external service credentials
How to Get Configuration
1. New Employee Onboarding
New developers joining the project, please follow this process to get configurations:
- Contact operations lead (see contact information at the end of document)
- State your needs:
- Your name and role
- Environment needed (development/testing)
- Specific services to access
- Receive configuration: Operations lead will provide configuration files or environment variables
- Local configuration: Place configuration information in project's
config.jsonor.envfile (Note: these files are in.gitignore, won't be committed to repository)
2. Configuration File Location
# Configuration files in project root (do not commit to git)
config.json # Main configuration file
.env # Environment variable configuration
env.template # Configuration template (reference, need to fill in real values)
3. Environment Variable Examples
Reference env.template file, your .env file typically contains the following types of configuration:
# MongoDB
MONGODB_URI=mongodb://...
MONGODB_DATABASE=...
# Redis
REDIS_HOST=...
REDIS_PORT=...
REDIS_PASSWORD=...
# Kafka
KAFKA_BOOTSTRAP_SERVERS=...
# ElasticSearch
ES_HOST=...
ES_PORT=...
Configuration Management Notes
⚠️ Security Standards
Prohibit committing sensitive configuration
- All configuration files containing passwords, keys, tokens must not be committed to git
- Check
.gitignoreincludes configuration files before committing - Using pre-commit hook can help detect sensitive information
Configuration file permissions
- Local configuration files should have appropriate permissions (only current user readable)
- Do not paste configuration content in public places (like chat records, documents)
Configuration update notifications
- If configuration is updated, operations team will notify relevant developers
- Update local configuration promptly after receiving notification
🔄 Configuration Change Process
If you need to:
- Add new configuration items
- Modify configuration structure
- Add new environments or services
Recommended process:
- Discuss with development lead: Confirm necessity and impact scope of configuration changes
- Contact operations lead: Explain configuration needs and reasons for changes
- Update configuration template: Update
env.templateand related documentation - Team notification: Notify all developers to sync update local configuration
Different Environment Description
| Environment | Purpose | Configuration Source | Notes |
|---|---|---|---|
| Development | Local development and debugging | Provided by operations | Usually connects to development database, data can be freely tested |
| Testing | Integration and functional testing | Auto-deployed configuration | Connects to test database, data periodically reset |
| Production | Live running services | Strictly controlled by operations | Only operations and authorized personnel can access |
Note: Developers usually only need development environment configuration. Testing and production environment configurations are managed by CI/CD and operations team.
🎨 Code Style Standards
Pre-commit Hook Configuration
The project uses pre-commit to unify code style. It's recommended to install pre-commit hook after first cloning the project.
Installation Steps
# One-click dev environment setup (sync deps + install hooks)
make dev-setup
Tip:
make dev-setupautomatically runsuv sync --devand installs pre-commit hooks. If you only need to install hooks separately, runmake setup-hooks.
Functions
Pre-commit hook will automatically execute the following checks before each commit:
- Code formatting: Format Python code using black/ruff
- Import sorting: Sort import statements using isort
- Code checking: Code quality check using ruff/flake8
- Type checking: Type check using pyright/mypy
- YAML/JSON format: Check configuration file format
- Trailing whitespace: Remove extra whitespace at end of files
Manual Check
# Run check on all files
pre-commit run --all-files
# Run check on staged files
pre-commit run
💬 Comment Standards
Core Principle
💡 Important Note: Always add sufficient comments
Good comments help team members quickly understand code intent, improving maintainability and Code Review efficiency.
Comment Requirements
1. Function-level Comments (Google-style Docstring)
Every function/method should have a clear Google-style docstring explaining:
- Description: What the function does
- Args: Type and purpose of each parameter
- Returns: Return value type and meaning
- Raises: Exceptions that may be thrown (if applicable)
# ✅ Recommended: Complete function-level comments
async def fetch_user_memories(
user_id: str,
limit: int = 100,
include_archived: bool = False
) -> list[Memory]:
"""
Fetch user's memory list.
Args:
user_id: User unique identifier
limit: Maximum number of memories to return, default 100
include_archived: Whether to include archived memories, default False
Returns:
User's memory list, sorted by creation time in descending order
Raises:
UserNotFoundError: When user does not exist
"""
...
2. Step-level Comments
In complex business logic, add comments at key steps to explain the purpose of each step:
# ✅ Recommended: Add comments at key steps
async def process_memory_extraction(raw_data: dict) -> Memory:
# 1. Validate input data integrity
validated_data = validate_input(raw_data)
# 2. Extract key information (people, events, time, etc.)
extracted_info = await extract_key_information(validated_data)
# 3. Generate vector embedding for subsequent retrieval
embedding = await generate_embedding(extracted_info.content)
# 4. Build memory object and persist
memory = Memory(
content=extracted_info.content,
embedding=embedding,
metadata=extracted_info.metadata
)
return memory
3. Complex Logic Explanation
For complex algorithms, business rules, or non-intuitive code, add detailed explanations:
# ✅ Recommended: Explain complex business rules
def calculate_memory_score(memory: Memory, query: str) -> float:
"""Calculate relevance score between memory and query"""
# Base similarity score (cosine similarity)
base_score = cosine_similarity(memory.embedding, query_embedding)
# Time decay factor: newer memories have higher weight
# Using exponential decay with half-life of 30 days
days_old = (now - memory.created_at).days
time_decay = math.exp(-0.693 * days_old / 30)
# Importance weighting: memories marked as important get 50% boost
importance_boost = 1.5 if memory.is_important else 1.0
return base_score * time_decay * importance_boost
Comment Style
- Use Chinese or English consistently within the same project/module
- Comments should be concise and clear, avoid redundancy
- Keep comments updated when code changes
- Don't comment obvious code
# ❌ Not recommended: Redundant comment
i = i + 1 # increment i by 1
# ✅ Recommended: Explain "why" not "what"
i = i + 1 # Skip header row, start processing from data rows
Checklist
Before submitting code, confirm:
- All public functions/methods have docstrings
- Complex business logic has step-level comments
- Non-intuitive code has explanatory comments
- Comments are in sync with code, no outdated comments
- Reviewers can quickly understand code intent
📖 API Specification Sync
Core Principle
💡 Important Note: Must synchronize API documentation when modifying API interfaces
API documentation is the key basis for frontend-backend collaboration and service integration. Inconsistency between documentation and actual API leads to integration issues and debugging difficulties.
Sync Requirements
When modifying API interfaces, must complete the following sync operations:
1. Update API Documentation Comments
Ensure code API documentation comments match actual behavior:
# ✅ Recommended: Keep documentation comments consistent with actual API
from fastapi import APIRouter, Query
router = APIRouter()
@router.get("/memories/{memory_id}")
async def get_memory(
memory_id: str,
include_embedding: bool = Query(False, description="Whether to return vector embedding")
) -> MemoryResponse:
"""
Get detailed information of specified memory.
- **memory_id**: Memory unique identifier
- **include_embedding**: Whether to include vector embedding data in response
Returns:
MemoryResponse: Memory details including content, metadata, etc.
Raises:
404: Memory not found
403: No permission to access this memory
"""
...
2. Update Schema Definition Files
If API request/response structure changes, update related schema definitions:
# Update Pydantic model
class MemoryResponse(BaseModel):
"""Memory response model"""
id: str = Field(..., description="Memory unique identifier")
content: str = Field(..., description="Memory content")
created_at: datetime = Field(..., description="Creation time")
# When adding new fields, add clear descriptions
embedding: list[float] | None = Field(None, description="Vector embedding, only returned on request")
3. Regenerate API Documentation Files
If the project uses auto-generated API documentation (e.g., OpenAPI/Swagger), ensure regeneration:
# Example: Regenerate OpenAPI documentation
python scripts/generate_openapi.py
# Or ensure FastAPI auto-generated docs are up to date
# Visit /docs or /redoc to verify
4. Notify Stakeholders
If it's a major API change, notify frontend and other dependent service developers.
Checklist
Before submitting API changes, confirm:
- API documentation comments updated and consistent with actual behavior
- Schema definition files (Pydantic models, etc.) updated
- Auto-generated API documentation files regenerated
- Frontend and other services can develop based on latest API specification
- If breaking changes, stakeholders have been notified
📄 Documentation Standards
Core Principle
💡 Important Note: Do not over-generate documentation
Documentation is an important supplement to code, but excessive documentation increases maintenance burden. Follow the "necessary and sufficient" principle.
When Documentation is Needed
| Scenario | Need Documentation | Notes |
|---|---|---|
| Small bug fix | ❌ No | Just explain in code comments |
| Small feature optimization | ❌ No | Explain in commit message and code comments |
| New API endpoint | ⚠️ Depends | API doc comments required, separate doc depends on complexity |
| New module/component | ✅ Yes | Write module introduction documentation |
| Large-scale refactoring | ✅ Yes | Document reasons, approach and impact |
| Architecture design changes | ✅ Yes | Document design decisions and architecture description |
| Complex business processes | ✅ Yes | Write process documentation |
Documentation Format Requirements
- Format: Use Markdown (
.md) format - Syntax: Follow standard Markdown syntax
Documentation Location
project_root/
├── docs/ # Documentation root
│ ├── api_docs/ # API documentation
│ │ └── memory_api.md
│ ├── dev_docs/ # Development documentation
│ │ └── development_standards.md
│ ├── architecture/ # Architecture documentation
│ │ └── system_design.md
│ └── guides/ # User guides
│ └── getting_started.md
Naming Convention
- Format:
{category}/{filename}.md - Examples:
api_docs/document_slice_api.mddev_docs/coding_standards.mdarchitecture/memory_system_design.md
Documentation Content Suggestions
A good document typically contains:
- Title and introduction: Explain the purpose of the document
- Background/motivation: Why this feature/change is needed
- Core content: Detailed explanation
- Examples: Code examples or usage examples
- Related documentation: Links to other related documents
Checklist
Before writing documentation, ask yourself:
- Is this change complex enough to need separate documentation?
- Are code comments already sufficient to explain the issue?
- Is the documentation in the correct directory?
- Is the documentation name clear and understandable?
⚡️ Async Programming Standards
Full Async Architecture Principles
The project adopts full async architecture, based on the following principles:
1. Single Event Loop Principle
- The entire application uses one main Event Loop
- Avoid creating new Event Loops in code (
asyncio.new_event_loop()) - Avoid using
asyncio.run()to start new loops in async context
2. About Using Threads and Processes ⚠️
💡 Important Note: Be cautious with multithreading and multiprocessing
The project is based on single Event Loop full async architecture, avoid the following operations:
# ❌ Not recommended: Creating threads
import threading
thread = threading.Thread(target=some_function)
thread.start()
# ❌ Not recommended: Using thread pool (unless special cases)
from concurrent.futures import ThreadPoolExecutor
executor = ThreadPoolExecutor()
# ❌ Not recommended: Creating processes
import multiprocessing
process = multiprocessing.Process(target=some_function)
process.start()
# ❌ Not recommended: Using process pool
from concurrent.futures import ProcessPoolExecutor
executor = ProcessPoolExecutor()
Why not recommended?
- May break single Event Loop architecture, causing concurrency issues
- Thread safety issues are complex, easy to introduce race conditions
- Resource management is difficult, may cause resource leaks
- May affect async context (contextvars) normal operation
- Debugging becomes harder, stack traces are complex
Special Case Handling
If you really need to use threads or processes (e.g., CPU-intensive computation, calling third-party libraries that don't support async), it's recommended to:
- Discuss with development lead in advance
- Explain why async solution cannot meet the needs
- Provide resource management plan (ensure threads/processes are properly closed)
- Go through Code Review
Allowed scenario examples:
# ✅ Special case: Calling sync libraries that don't support async (after discussion)
import asyncio
from concurrent.futures import ThreadPoolExecutor
# Globally shared thread pool, limit max threads
_EXECUTOR = ThreadPoolExecutor(max_workers=4)
async def call_sync_library(data):
"""Call third-party library that doesn't support async (confirmed with lead)"""
loop = asyncio.get_event_loop()
# Run in thread pool to avoid blocking main loop
result = await loop.run_in_executor(
_EXECUTOR,
sync_blocking_function,
data
)
return result
3. Async Function Definition
I/O operations should use async functions:
# ✅ Correct: Async function
async def fetch_user_data(user_id: str) -> dict:
async with httpx.AsyncClient() as client:
response = await client.get(f"/users/{user_id}")
return response.json()
# ❌ Wrong: Sync I/O
def fetch_user_data(user_id: str) -> dict:
response = requests.get(f"/users/{user_id}")
return response.json()
4. Database Operations
# ✅ Correct: Using async database driver
from pymongo import AsyncMongoClient
async def get_user(db, user_id: str):
return await db.users.find_one({"_id": user_id})
# ❌ Wrong: Using sync driver
from pymongo import MongoClient
def get_user(db, user_id: str):
return db.users.find_one({"_id": user_id})
5. HTTP Client
# ✅ Correct: Using httpx.AsyncClient
import httpx
async def call_api(url: str):
async with httpx.AsyncClient() as client:
response = await client.get(url)
return response.json()
# ❌ Wrong: Using requests
import requests
def call_api(url: str):
response = requests.get(url)
return response.json()
6. Concurrent Processing
Use asyncio.gather() for concurrent operations:
# ✅ Correct: Execute multiple tasks concurrently
async def fetch_multiple_users(user_ids: list[str]):
tasks = [fetch_user_data(uid) for uid in user_ids]
results = await asyncio.gather(*tasks)
return results
# ❌ Wrong: Serial execution
async def fetch_multiple_users(user_ids: list[str]):
results = []
for uid in user_ids:
result = await fetch_user_data(uid)
results.append(result)
return results
7. Prohibit I/O Operations in for Loops ⚠️
💡 Important Note: Avoid serial I/O operations in loops
Doing database access, API calls and other I/O operations in for loops causes serious performance issues, because each operation needs to wait for the previous one to complete, unable to take advantage of async concurrency.
❌ Wrong example: I/O operations in loops
# Wrong: Serial database access in loop
async def get_users_info(user_ids: list[str]):
results = []
for user_id in user_ids:
# Each loop iteration waits for database return, very poor performance
user = await db.users.find_one({"_id": user_id})
results.append(user)
return results
# Wrong: Serial API calls in loop
async def fetch_user_profiles(user_ids: list[str]):
profiles = []
for user_id in user_ids:
# Each loop iteration waits for API response, wasting time
response = await api_client.get(f"/users/{user_id}")
profiles.append(response.json())
return profiles
# Wrong: Batch database inserts in loop
async def save_messages(messages: list[dict]):
for msg in messages:
# Each message inserted separately, very inefficient
await db.messages.insert_one(msg)
✅ Correct example: Using concurrent or batch operations
# Correct: Using asyncio.gather for concurrent execution
async def get_users_info(user_ids: list[str]):
tasks = [db.users.find_one({"_id": uid}) for uid in user_ids]
results = await asyncio.gather(*tasks)
return results
# Correct: Using asyncio.gather for concurrent API calls
async def fetch_user_profiles(user_ids: list[str]):
tasks = [api_client.get(f"/users/{uid}") for uid in user_ids]
responses = await asyncio.gather(*tasks)
return [r.json() for r in responses]
# Correct: Using batch insert operation
async def save_messages(messages: list[dict]):
if messages:
await db.messages.insert_many(messages)
# Correct: Using database's in query instead of loop query
async def get_users_info(user_ids: list[str]):
# Single query to get all data
cursor = db.users.find({"_id": {"$in": user_ids}})
results = await cursor.to_list(length=None)
return results
Performance Comparison
Assuming 100 users, each database query takes 10ms:
- ❌ Loop serial query: 100 × 10ms = 1000ms (1 second)
- ✅ Concurrent query: ~10ms (almost simultaneous completion)
- ✅ Batch query: ~10ms (single query)
Exception Cases
In rare cases you may need to do I/O in loops, but must meet the following conditions:
- Subsequent operations depend on previous result: Must wait for previous operation to complete before next one
- Rate limiting needs: Need to control concurrency to avoid pressure on external services
- Approved by development lead
# Allowed: Serial operations with dependencies (comment explaining reason)
async def process_workflow(steps: list[dict]):
result = None
for step in steps:
# Each step depends on previous step's result, cannot be concurrent
result = await execute_step(step, previous_result=result)
return result
# Allowed: Using semaphore to control concurrency (comment explaining reason)
async def fetch_with_rate_limit(urls: list[str]):
# Limit max 5 concurrent requests to avoid triggering external API rate limiting
semaphore = asyncio.Semaphore(5)
async def fetch_one(url: str):
async with semaphore:
return await api_client.get(url)
tasks = [fetch_one(url) for url in urls]
return await asyncio.gather(*tasks)
🕐 Timezone Awareness Standards
Core Principle
💡 Important Note: All time fields must be timezone-aware
When handling date and time data, must ensure all time fields carry timezone information to avoid data errors and business issues caused by unclear timezone.
⚠️ Prohibit direct use of datetime module standard methods
The project uniformly uses utility functions from common_utils/datetime_utils.py for time handling, prohibit direct use of:
- ❌
datetime.datetime.now() - ❌
datetime.datetime.utcnow() - ❌
datetime.datetime.today()
Must use project-provided utility functions:
- ✅
get_now_with_timezone()- Get current time (with timezone) - ✅
from_timestamp()- Convert from timestamp - ✅
from_iso_format()- Convert from ISO format string - ✅
to_iso_format()- Convert to ISO format string - ✅
to_timestamp()/to_timestamp_ms()- Convert to timestamp
Timezone Handling Rules
1. Input Data Timezone Requirements
All time fields entering the system must meet:
- Must carry timezone info: All datetime type fields must be timezone-aware
- Default timezone: If input data doesn't have timezone info, treat it as the timezone configured by environment variable
TZ(default UTC) - Storage format: When storing in database, recommend converting to UTC timezone uniformly, but must preserve timezone info
2. Python Implementation Standards
✅ Correct example: Using project utility functions
from common_utils.datetime_utils import (
get_now_with_timezone,
from_timestamp,
from_iso_format,
to_iso_format,
to_timestamp_ms,
to_timezone
)
# Method 1: Get current time (automatically with timezone, configured by TZ env var, default UTC)
now = get_now_with_timezone()
# Returns: datetime.datetime(2025, 9, 16, 12, 17, 41, tzinfo=zoneinfo.ZoneInfo(key='UTC'))
# Method 2: Convert from timestamp (auto-detect seconds/milliseconds, auto-add timezone)
dt = from_timestamp(1758025061)
dt_ms = from_timestamp(1758025061000)
# Method 3: Convert from ISO string (auto-handle timezone)
dt = from_iso_format("2025-09-15T13:11:15.588000") # No timezone, auto-add default timezone
dt_with_tz = from_iso_format("2025-09-15T13:11:15Z") # Has timezone, preserve original then convert
# Method 4: Format to ISO string (auto-include timezone)
iso_str = to_iso_format(now)
# Returns: "2025-09-16T12:20:06.517301Z"
# Method 5: Convert to timestamp
ts = to_timestamp_ms(now)
# Returns: 1758025061123
❌ Wrong example: Direct use of datetime module
import datetime
# ❌ Wrong: Prohibit using datetime.datetime.now()
naive_dt = datetime.datetime.now() # Timezone unclear, prohibit!
# ❌ Wrong: Prohibit using datetime.datetime.utcnow()
dt = datetime.datetime.utcnow() # Deprecated in Python 3.12+, prohibit!
# ❌ Wrong: Prohibit using datetime.datetime.today()
dt = datetime.datetime.today() # Timezone unclear, prohibit!
# ❌ Wrong: Manually creating naive datetime
naive_dt = datetime.datetime(2025, 1, 1, 12, 0, 0) # No timezone info
🔧 How to fix existing code
# Old code (wrong)
import datetime
now = datetime.datetime.now()
# New code (correct)
from common_utils.datetime_utils import get_now_with_timezone
now = get_now_with_timezone()
# ----------------
# Old code (wrong)
from datetime import datetime
dt = datetime(2025, 1, 1, 12, 0, 0)
# New code (correct)
from common_utils.datetime_utils import from_iso_format
dt = from_iso_format("2025-01-01T12:00:00") # Auto-add default timezone
# ----------------
# Old code (wrong)
ts = int(datetime.now().timestamp() * 1000)
# New code (correct)
from common_utils.datetime_utils import get_now_with_timezone, to_timestamp_ms
ts = to_timestamp_ms(get_now_with_timezone())
Checklist
During code review, please confirm:
- Prohibit direct use of
datetime.datetime.now(), must useget_now_with_timezone() - Prohibit direct use of
datetime.datetime.utcnow()ordatetime.datetime.today() - All time retrieval goes through utility functions in
common_utils/datetime_utils.py - Time parsed from external input uses
from_iso_format()orfrom_timestamp() - Time formatting uses
to_iso_format()instead of manually calling.isoformat() - Timestamp conversion uses
to_timestamp_ms()instead of manual calculation - Database schema uses timezone-aware types (e.g.,
timestamptz) - API response time strings include timezone info (ISO 8601 format)
- Test data in unit tests all have timezone info
🏛️ Data Access Standards
Core Principle
💡 Important Note: All external storage access must go through infra layer repository
When handling databases, search engines and other external storage systems, must follow strict layered architecture principles. All data read/write operations must be converged to infra_layer repository layer, prohibit direct calls to external storage capabilities in business layer or other upper layers.
⚠️ Prohibit direct external storage access in these layers
- ❌
biz_layer(Business layer) - ❌
memory_layer(Memory layer) - ❌
agentic_layer(Agent layer) - ❌ API interface layer (
api_specs) - ❌ Application layer (
app.py, controllers, etc.)
✅ Must access through
infra_layer/adapters/out/persistence/repository/- Database accessinfra_layer/adapters/out/search/repository/- Search engine access
Why This Standard?
1. Separation of Concerns
Following Hexagonal Architecture and Clean Architecture principles:
- Business layer: Focus on business logic, don't care where data comes from
- Infrastructure layer: Handle all external system interaction details
- Isolate changes: When changing database or search engine, only need to modify infra layer
2. Testability
# ✅ Benefit: Business layer depends on abstract interface, easy to mock test
async def process_user_memory(user_id: str, memory_repo: MemoryRepository):
"""Business logic doesn't depend on specific implementation"""
memories = await memory_repo.find_by_user_id(user_id)
# Business processing...
# Can easily replace with mock during testing
mock_repo = MockMemoryRepository()
await process_user_memory("user_1", mock_repo)
3. Code Reuse and Consistency
- Avoid repeatedly writing same database query logic in multiple places
- Unified exception handling, logging, performance monitoring
- Unified data transformation, validation
4. Centralized Performance Optimization
- Index optimization, query optimization implemented uniformly in repository layer
- Cache strategies managed uniformly
- Batch operation optimization done in one place, benefits entire project
Correct Architecture Layering
┌─────────────────────────────────────────┐
│ API Layer (api_specs, app.py) │
│ - Receive requests, return responses │
└─────────────┬───────────────────────────┘
│ calls
▼
┌─────────────────────────────────────────┐
│ Business Layer (biz_layer) │
│ - Business logic processing │
│ - Depends on abstract interfaces (Port)│
└─────────────┬───────────────────────────┘
│ dependency injection
▼
┌─────────────────────────────────────────┐
│ Memory Layer (memory_layer) │
│ - Memory management logic │
│ - Depends on abstract interfaces (Port)│
└─────────────┬───────────────────────────┘
│ dependency injection
▼
┌─────────────────────────────────────────┐
│ Infrastructure Layer (infra_layer) │
│ - Repository implementation (Adapter) │
│ - Directly operate database/search │
│ - MongoDB, PostgreSQL, ES, Milvus │
└─────────────────────────────────────────┘
Checklist
When writing or reviewing code, please confirm:
- Are database operations in infra_layer/repository?
- Are search engine operations in infra_layer/repository?
- Does business layer depend on abstract interfaces (Port) not concrete implementations?
- Is dependency injection used to pass repository?
- Avoid directly creating database connections in business/API/application layers?
- Avoid directly using MongoDB/PostgreSQL/ES/Milvus clients in business layer?
- Has new Repository been registered in dependency injection container?
- Do Repository methods have clear business semantics (not exposing underlying implementation details)?
📥 Import Standards
PYTHONPATH Management
💡 Important Note: PYTHONPATH needs unified management
The project uniformly manages PYTHONPATH and module import paths. Changes involving path configuration should be discussed with development lead before unified configuration.
Why Unified Management?
- Chaotic import paths may cause modules not found or import errors
- Inconsistent paths across environments (dev/test/prod) may cause deployment issues
- Inconsistent IDE configuration may affect team collaboration
- Mixing relative and absolute imports increases code maintenance difficulty
Management Scope
The following directories in the project should maintain unified import paths:
src/: Main business codetests/: Test codeunit_test/: Unit testsevaluation/: Evaluation scripts- Other directories needing to be imported (e.g.,
demo/)
Recommended Practices
Unified project root directory
- Project root is
/Users/admin/memsys(or corresponding deployment path) - src directory added to PYTHONPATH, import directly from module name
- Project root is
Import standard examples
# ✅ Recommended: Absolute import (src already in PYTHONPATH)
from core.memory.manager import MemoryManager
from infra_layer.adapters.out.db import MongoDBAdapter
from tests.fixtures.mock_data import get_mock_user
# ✅ Recommended: Import in test files
from unit_test.email_data_constructor import construct_email
# ❌ Not recommended: Cross-level relative import
from ...core.memory.manager import MemoryManager
# ❌ Not recommended: Including src prefix (src already in PYTHONPATH, no prefix needed)
from src.core.memory.manager import MemoryManager
# ❌ Not recommended: Using sys.path.append to temporarily modify path
import sys
sys.path.append("../src") # May cause environment inconsistency
Prefer Absolute Imports
💡 Important Note: Recommend absolute imports, avoid relative imports
Why Recommend Absolute Imports?
Although relative imports are more concise in some scenarios, they have these issues:
- Poor readability:
from ...core.memory import Manageris less intuitive thanfrom core.memory import Manager - Difficult refactoring: Moving files requires modifying all relative import levels
- Complex debugging: Stack traces with relative import paths are unclear
- Tool support: IDE and static analysis tools support absolute imports better
- **Te
…(truncated)