Refactor score_growth_rates.py for Improved Modularity and Testability
Overview
Refactor scripts/score_growth_rates.py to follow the same modular architecture successfully implemented in the score_models.py refactor. This will improve code maintainability, testability, and reusability while following test-driven development principles.
Motivation
The current score_growth_rates.py script contains ~415 lines of monolithic code with:
- Hardcoded configuration constants
- Complex nested logic in a single large function
- Mixed concerns (loading, filtering, processing, exporting)
- Difficult to test individual components
- Limited reusability of functionality
Proposed Changes
1. Create antigentools/growth_scoring/ Module
Following the pattern from antigentools/scoring/, create a new module with:
Module Structure
antigentools/growth_scoring/
├── __init__.py # Public API exports
├── config.py # Configuration management
├── loaders.py # Data loading functions
├── filters.py # Filtering logic
├── collectors.py # Result collection dataclasses
├── exporters.py # Export functionality
└── processors.py # Main processing pipeline
2. Configuration Management (config.py)
Current State: Hardcoded constants (lines 51-58)
CONNECT_GAPS = True
MIN_SEGMENT_LENGTH = 3
MIN_SEQUENCE_COUNT = 10
MIN_VARIANT_FREQUENCY = 0.01
EPSILON = 1e-3
MIN_TOTAL_SEQUENCES = 300
CONVERGENCE_THRESHOLD = 0.5
Proposed:
- Create
GrowthRateConfig dataclass with validation
- Create
ConvergenceConfig dataclass for VI diagnostics settings
- Support loading from YAML configuration files
- Provide sensible defaults with validation
3. Data Loaders (loaders.py)
Current State: Mixed into main processing function
- RT file discovery (lines 186-187)
- Growth rate loading (lines 217-228)
- Convergence diagnostics loading (lines 253-296)
Proposed:
discover_rt_files(build: str, models: List[str]) -> List[RTFile]
load_growth_rates(build: str, model: str, location: str, pivot_date: str) -> pd.DataFrame
load_convergence_diagnostics(path: Path) -> Dict[str, Any]
4. Filters (filters.py)
Current State: Inline filtering logic
- Window sequence count filtering (lines 210-216)
Proposed:
filter_by_sequence_count(df: pd.DataFrame, min_count: int) -> pd.DataFrame
filter_windows_by_total_sequences(windows: List[Window], min_total: int) -> List[Window]
- Composable filter functions following functional programming patterns
5. Result Collectors (collectors.py)
Current State: Three large dictionaries (lines 122-178)
results_dict = {'pivot_date': [], 'model': [], ...}
variant_results_dict = {'variant': [], 'pivot_date': [], ...}
diagnostics_dict = {'pivot_date': [], 'model': [], ...}
Proposed Dataclasses:
WindowResults: Type-safe container for window-level metrics
VariantResults: Type-safe container for variant-level metrics
ConvergenceDiagnostics: Type-safe container for VI diagnostics
ResultsCollector: Aggregates all results with helper methods
6. Exporters (exporters.py)
Current State:
export_growth_rates_data() function (lines 60-96)
- Inline CSV saving (lines 385-387)
Proposed:
export_growth_rates(df: pd.DataFrame, output_dir: Path, metadata: ExportMetadata)
export_window_results(results: WindowResults, path: Path)
export_variant_results(results: VariantResults, path: Path)
export_diagnostics(diagnostics: ConvergenceDiagnostics, path: Path)
7. Processors (processors.py)
Current State: Monolithic process_all_model_results() (lines 99-393)
Proposed:
process_growth_rate_window(window: RTWindow, config: GrowthRateConfig) -> WindowResult
process_all_windows(windows: List[RTWindow], config: GrowthRateConfig) -> ResultsCollector
- Clear separation of orchestration from individual processing steps
Testing Strategy
Following TDD principles, tests will be written before implementation:
Test Structure
tests/test_growth_scoring/
├── test_config.py # Configuration validation tests
├── test_loaders.py # Data loading tests
├── test_filters.py # Filter function tests
├── test_collectors.py # Dataclass and collection tests
├── test_exporters.py # Export functionality tests
├── test_processors.py # Pipeline integration tests
└── fixtures/ # Test data fixtures
Key Test Cases
- Configuration: Validation, defaults, YAML loading
- Loaders: File discovery, data loading, error handling
- Filters: Edge cases, empty data, composition
- Collectors: Data integrity, type safety, aggregation
- Exporters: File creation, format correctness
- Processors: End-to-end pipeline, error recovery
Benefits
- Testability: Each component can be tested in isolation
- Maintainability: Clear separation of concerns
- Reusability: Functions can be used in other contexts
- Type Safety: Dataclasses provide runtime validation
- Configuration: Flexible configuration without code changes
- Extensibility: Easy to add new filters, metrics, or export formats
Implementation Plan
- Create module structure and placeholder files
- Write comprehensive tests for each module
- Implement modules following TDD (red-green-refactor)
- Refactor main script to use new modules
- Ensure backward compatibility and identical outputs
- Update documentation and examples
Backward Compatibility
- The refactored script will maintain the same CLI interface
- Output files will have identical format and content
- Existing configuration files will continue to work
Refactor score_growth_rates.py for Improved Modularity and Testability
Overview
Refactor
scripts/score_growth_rates.pyto follow the same modular architecture successfully implemented in thescore_models.pyrefactor. This will improve code maintainability, testability, and reusability while following test-driven development principles.Motivation
The current
score_growth_rates.pyscript contains ~415 lines of monolithic code with:Proposed Changes
1. Create
antigentools/growth_scoring/ModuleFollowing the pattern from
antigentools/scoring/, create a new module with:Module Structure
2. Configuration Management (
config.py)Current State: Hardcoded constants (lines 51-58)
Proposed:
GrowthRateConfigdataclass with validationConvergenceConfigdataclass for VI diagnostics settings3. Data Loaders (
loaders.py)Current State: Mixed into main processing function
Proposed:
discover_rt_files(build: str, models: List[str]) -> List[RTFile]load_growth_rates(build: str, model: str, location: str, pivot_date: str) -> pd.DataFrameload_convergence_diagnostics(path: Path) -> Dict[str, Any]4. Filters (
filters.py)Current State: Inline filtering logic
Proposed:
filter_by_sequence_count(df: pd.DataFrame, min_count: int) -> pd.DataFramefilter_windows_by_total_sequences(windows: List[Window], min_total: int) -> List[Window]5. Result Collectors (
collectors.py)Current State: Three large dictionaries (lines 122-178)
Proposed Dataclasses:
WindowResults: Type-safe container for window-level metricsVariantResults: Type-safe container for variant-level metricsConvergenceDiagnostics: Type-safe container for VI diagnosticsResultsCollector: Aggregates all results with helper methods6. Exporters (
exporters.py)Current State:
export_growth_rates_data()function (lines 60-96)Proposed:
export_growth_rates(df: pd.DataFrame, output_dir: Path, metadata: ExportMetadata)export_window_results(results: WindowResults, path: Path)export_variant_results(results: VariantResults, path: Path)export_diagnostics(diagnostics: ConvergenceDiagnostics, path: Path)7. Processors (
processors.py)Current State: Monolithic
process_all_model_results()(lines 99-393)Proposed:
process_growth_rate_window(window: RTWindow, config: GrowthRateConfig) -> WindowResultprocess_all_windows(windows: List[RTWindow], config: GrowthRateConfig) -> ResultsCollectorTesting Strategy
Following TDD principles, tests will be written before implementation:
Test Structure
Key Test Cases
Benefits
Implementation Plan
Backward Compatibility