# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

This is a Machine Learning system for predicting medical claim denials ("glosas") using XGBoost models. The system uses specialized models trained for different types of medical guides (Internação, SADT, Consulta, Honorários, Odonto) with Out-of-Vocabulary (OOV) detection for unknown procedure codes.

### System Requirements
- **Docker Desktop**: Must be running for containerized deployment
- **Python**: 3.8+ for local development
- **PHP**: 7.4+ for web interface (local development only)
- **Ports**: 8010 (API), 8011 (Web interface)

## Core Architecture

### Model Strategy
- **Type-Specific Models**: One specialized model per medical guide type (tipo_guia 1-5)
- **OOV Detection**: Models with `_oov.pkl` suffix include Out-of-Vocabulary detection to penalize unknown procedure codes
- **Model Priority**: OOV models are always prioritized over standard models when both exist
- **Factoring Mode**: Ultra-conservative mode with fixed 0.15 threshold for receivables evaluation

### Key Components

1. **Training Pipeline** (`treina_por_tipo_guia_oov.py`)
   - Trains separate models per guide type
   - Generates synthetic OOV examples to teach models about unknown codes
   - Creates features: `codigo_desconhecido`, `codigo_suspeito`, `frequencia_codigo`, `categoria_codigo`
   - Saves models as `modelo_*_tipo{N}_oov.pkl`

2. **Prediction API** (`api_predicao.py`)
   - `PreditorAPI` class handles all prediction logic
   - Automatic model selection based on guide type
   - OOV penalty system: unknown codes receive high risk scores (0.8-1.0)
   - Factoring mode applies additional penalties for financial risk assessment

3. **REST API Server** (`api_server.py`)
   - Flask server on port 8010
   - Endpoints: `/`, `/health`, `/modelos`, `/predizer`, `/predizer/batch`
   - CORS enabled for development

4. **Smart Predictor** (`predizer_inteligente.py`)
   - Batch file processing with automatic model selection
   - Statistics and distribution analysis

5. **Data Extraction Pipeline** (`extrai_dados_e_treina.py`)
   - Connects to MySQL database using .env credentials
   - Extracts training data directly from source
   - Automatically trains models after extraction

## Development Commands

### Quick Start
```bash
# Fastest way to start everything (Docker Desktop must be running)
./start.sh

# This script will:
# - Check Docker is running
# - Stop previous services
# - Build and start containers (API + Web)
# - Wait for services to be ready
# - Display access URLs
```

### Environment Setup
```bash
# Always use virtual environment for local development
python3 -m venv venv
source venv/bin/activate  # On macOS/Linux
# venv\Scripts\activate  # On Windows
pip install -r requirements.txt
```

### Web Interface for Model Management
```bash
# With Docker Compose (recommended - included automatically)
docker-compose up -d
# Access: http://localhost:8011

# Or start PHP server locally for development
cd web/
php -S localhost:8000
# Access: http://localhost:8000

# Features:
# - Upload models (.pkl files up to 100MB)
# - View all models with metadata
# - Delete old models
# - Real-time API status
# - Auto-detection of OOV models and guide types
```

### Training Models
```bash
# Option 1: Extract data from MySQL and train (full pipeline)
python extrai_dados_e_treina.py
# Requires .env file with database credentials

# Option 2: Train from existing CSV with OOV detection (recommended)
python treina_por_tipo_guia_oov.py 20251112_treinamento.csv

# Train specific type only
python treina_por_tipo_guia_oov.py 20251112_treinamento.csv 3  # Consulta

# Models are saved to root directory by default
# Move to modelos/ for production use
```

### Running Predictions
```bash
# Using intelligent predictor (analyzes file and selects best model)
python predizer_inteligente.py guias.csv

# Using JSON predictor
python predizer_json.py dados.json
```

### API Development
```bash
# Start API server (development)
python run_api.py
# or
./start_api.sh

# Using Docker Compose (production)
docker-compose up -d
docker-compose logs -f
docker-compose down

# Direct Docker
docker build -t predicao-glosas .
docker run -d -p 8010:8010 predicao-glosas
```

### Testing API
```bash
# Health check
curl http://localhost:8010/health

# List models
curl http://localhost:8010/modelos

# Single prediction
curl -X POST http://localhost:8010/predizer \
  -H "Content-Type: application/json" \
  -d '{"protocolo": 123, "tipo_guia": 3, "cod": 10101012, "quantidade": 1}'
```

## Important Implementation Details

### Data Preparation Flow
1. Convert date columns to datetime (removes after feature extraction)
2. Create temporal features: `dias_envio`, `dias_ate_cirurgia`, `mes_cirurgia`, `dia_semana_cirurgia`
3. For OOV models: add `codigo_desconhecido`, `codigo_suspeito`, `frequencia_codigo`, `categoria_codigo`
4. Encode categorical features with saved LabelEncoders
5. Align with model's expected feature_names
6. XGBoost cannot accept datetime columns - must be removed before prediction

### OOV Detection Logic
```python
# In api_predicao.py:
if codigo not in codigos_conhecidos:
    is_unknown = True
    suspicion_score = 0.8  # Base penalty
    if codigo > 88000000:
        suspicion_score = 1.0  # Maximum penalty for very suspicious codes
    probabilidade_ajustada = max(probabilidade, penalty)
```

### Factoring Mode Specifics
- Fixed threshold: 0.15 (15% risk)
- Unknown codes: force rejection (adds 0.85 penalty → 1.0 total)
- Multiple procedures: adds 0.10 penalty
- Risk levels: <10% = low, 10-15% = medium, >15% = very_high
- Returns `factoring` object with approval recommendation

### Model File Naming Convention
- General: `modelo_*_YYYYMMDD_*.pkl`
- Type-specific: `modelo_*_tipo{1-5}.pkl`
- OOV-enabled: `modelo_*_tipo{1-5}_oov.pkl`
- Advanced: `modelo_consulta_avancado.pkl`

### Model Loading Strategy
- **Primary location**: `modelos/` directory (prioritized)
- **Fallback location**: Root directory `./`
- **Selection**: Always loads the most recent model (sorted by modification time)
- **Priority**: OOV models > Normal models
- **Dynamic loading**: No hardcoded paths - automatically finds latest models

### Critical Features by Type
- **All types**: tipo_guia, id_convenio, cod, quantidade
- **Internação (tipo 1)**: tipo_atendimento, via_acesso, dias_ate_cirurgia
- **SADT (tipo 2)**: quantidade, tipo_atendimento
- **Consulta (tipo 3)**: cod (procedure code is most important)
- **OOV models add**: codigo_desconhecido, codigo_suspeito, frequencia_codigo, categoria_codigo

## Code Patterns to Follow

### When Adding New Features
1. Add to `preparar_dados()` or `preparar_dados_com_oov()` in api_predicao.py
2. Update training script to include feature
3. Retrain models to get feature in feature_names list
4. Test with OOV and non-OOV models

### When Modifying Thresholds
- Normal models: threshold stored in `modelo_info['best_threshold']`
- OOV models: also have `modelo_info['oov_threshold']` for unknown codes
- Factoring mode: uses fixed `THRESHOLD_FACTORING = 0.15`

### Error Handling Pattern
```python
# Always handle missing features gracefully
for feat in feature_names:
    if feat in df.columns:
        X[feat] = df[feat]
    else:
        X[feat] = 0  # Default value for missing features
```

## File Organization

### Data Files
- Training CSVs: `YYYYMMDD_treinamento*.csv`
- Type-specific: `*_consulta.csv`, `*_internacao.csv`, `*_sadt.csv`

### Model Files
- **Primary**: `modelos/` directory - all production models stored here
- **Fallback**: Root directory `./` - checked if models not found in modelos/
- **Auto-selection**: API automatically loads most recent OOV models from modelos/

### Configuration
- `.env`: Database credentials (not versioned) - required for `extrai_dados_e_treina.py`
- `requirements.txt`: Main dependencies for local development
- `requirements-docker.txt`: Docker-specific dependencies
- `start.sh`: Automated startup script with health checks
- `start_api.sh`: Legacy script for starting API only

### Scripts
- `start.sh`: Full system startup (API + Web interface + health checks)
- `start_api.sh`: Start API only (legacy)
- `run_api.py`: Direct Python API startup

## Testing Strategy

1. **Model Loading**: Check `PreditorAPI.__init__()` finds all models
2. **OOV Detection**: Test with unknown codes (88000000-99999999 range)
3. **Factoring Mode**: Verify ultra-conservative behavior with `modo_factoring=True`
4. **Multi-procedure**: Test protocols with multiple items (same protocolo value)
5. **API Integration**: Use exemplos_api.md for curl/PHP/Python examples

## Common Issues and Solutions

### XGBoost Feature Errors
- Ensure datetime columns removed before prediction
- Check all categorical features are encoded
- Verify feature order matches model's feature_names

### OOV Model Not Loading
- Check filename ends with `_oov.pkl`
- Verify model dict contains `codigos_conhecidos` and `oov_features`
- Check model is in root directory, not subdirectory

### Docker Build Issues
- Ensure all `.pkl` files are in root directory
- Check `requirements-docker.txt` has all dependencies
- Verify ports 8010 (API) and 8011 (Web) are available
- Use `./start.sh` for automated startup with health checks

### Port Conflicts
```bash
# Kill process using API port
lsof -ti:8010 | xargs kill -9

# Kill process using Web port
lsof -ti:8011 | xargs kill -9

# Then restart
./start.sh
```

### Database Connection Issues
- Check `.env` file exists with correct credentials
- Required for `extrai_dados_e_treina.py` only
- Not needed for API or predictions with existing models

## API Response Structure

```json
{
  "protocolo": 123,
  "tipo_guia": "Consulta",
  "analise_procedimentos": [
    {
      "id_procedimento": "PROC-001",
      "cod": "10101012",
      "risco": 0.15,
      "codigo_desconhecido": true,  // Only if OOV model
      "alerta": "Código nunca visto no treinamento"
    }
  ],
  "resultado": {
    "media": 0.15,
    "maximo": 0.15,
    "risco": "baixo",  // baixo|medio|alto|muito_alto
    "predicao": "aprovado",  // aprovado|glosa_provavel
    "threshold": 0.77
  },
  "modelo_usado": {
    "arquivo": "modelo_20251112_treinamento_consulta_tipo3_oov.pkl",
    "versao": "v1",
    "tipo": "OOV",
    "tem_deteccao_oov": true
  },
  "modo_analise": "FACTORING_ULTRA_CONSERVADOR",  // Only if factoring mode
  "factoring": {  // Only if factoring mode
    "aprovado_para_antecipacao": true,
    "margem_seguranca": 0.35,
    "confianca": "ALTA",
    "recomendacao": {
      "decisao": "APROVAR",
      "motivo": "Risco 15.0% dentro da margem segura",
      "acao": "Pode aprovar para antecipação com segurança"
    }
  },
  "alertas": ["⚠️ 1 código(s) nunca visto(s) no treinamento"],
  "tempo_processamento_ms": 45.23
}
```

## Model Training Data Requirements

- Minimum rows per type: ~1000 for reliable model
- Required columns: protocolo, tipo_guia, id_convenio, cod, quantidade, glosado
- Optional but recommended: data_cirurgia, tipo_atendimento, via_acesso
- Target variable: `glosado` (0 or 1)
- OOV training adds 15% synthetic examples with unknown codes marked as glosado=1
