Stellium v0.22.0
Cookbooks/ Statistical Analysis

Statistical Analysis

Calculate hundreds of charts at once, convert them to pandas DataFrames, query them by astrological criteria, and run statistics across the whole set.

The analysis module provides tools for:

  • Batch chart calculation - Calculate 100s-1000s of charts efficiently

  • DataFrame conversion - Export charts to pandas DataFrames

  • Research queries - Filter charts by astrological criteria

  • Statistical analysis - Aggregate statistics across chart collections

  • Export utilities - Save to CSV, JSON, Parquet

Installation

The analysis module requires pandas (optional dependency):

pip install stellium[analysis]
# Imports
import pandas as pd

from stellium.analysis import (
    BatchCalculator,
    ChartQuery,
    ChartStats,
    aspects_to_dataframe,
    charts_to_dataframe,
    export_csv,
    export_json,
    positions_to_dataframe,
)
from stellium.core.native import Native
from stellium.engines.patterns import AspectPatternAnalyzer

1. BatchCalculator

Efficiently calculate many charts at once. BatchCalculator supports:

  • Loading from the NotableRegistry (with filters)

  • Loading from a list of Native objects

  • Generator-based calculation (memory efficient)

  • Progress callbacks

1.1 From NotableRegistry

Load charts from the built-in database of notable births and events.

# Calculate charts for all notables in the registry
all_charts = BatchCalculator.from_registry().calculate_all()
print(f"Calculated {len(all_charts)} charts")
Calculated 211 charts
# Filter by category
scientist_charts = BatchCalculator.from_registry(category="scientist").calculate_all()
print(f"Calculated {len(scientist_charts)} scientist charts")
Calculated 20 scientist charts
# Multiple filters: verified high-quality data only
verified_charts = BatchCalculator.from_registry(
    verified=True, data_quality="AA"
).calculate_all()
print(f"Calculated {len(verified_charts)} verified AA-quality charts")
Calculated 54 verified AA-quality charts

1.2 From Native Objects

Calculate charts from your own data.

# Create sample Native objects
sample_natives = [
    Native("2000-01-01 12:00", "New York, NY", name="Person A"),
    Native("1995-06-21 08:30", "Los Angeles, CA", name="Person B"),
    Native("1988-12-15 22:00", "Chicago, IL", name="Person C"),
    Native("1975-03-20 06:00", "Seattle, WA", name="Person D"),
    Native("1960-09-10 14:30", "Miami, FL", name="Person E"),
]

# Calculate charts
custom_charts = BatchCalculator.from_natives(sample_natives).calculate_all()
print(f"Calculated {len(custom_charts)} custom charts")
Calculated 5 custom charts

1.3 With Aspects and Analyzers

Configure the calculation with aspect detection and pattern analysis.

# Calculate with aspects and pattern detection
charts_with_aspects = (
    BatchCalculator.from_natives(sample_natives)
    .with_aspects()  # Enable aspect calculation
    .add_analyzer(AspectPatternAnalyzer())  # Detect Grand Trines, T-Squares, etc.
    .calculate_all()
)

# Check aspects for first chart
print(f"First chart has {len(charts_with_aspects[0].aspects)} aspects")
First chart has 40 aspects

1.4 Progress Tracking

Track progress for long-running batch calculations.

# Define a progress callback
def show_progress(current, total, name):
    print(f"  [{current}/{total}] Calculating: {name}")


# Calculate with progress tracking
print("Calculating charts with progress:")
charts = (
    BatchCalculator.from_natives(sample_natives[:3])  # Just first 3 for demo
    .with_progress(show_progress)
    .calculate_all()
)
Calculating charts with progress:
  [1/3] Calculating: Person A
  [2/3] Calculating: Person B
  [3/3] Calculating: Person C

1.5 Generator Mode (Memory Efficient)

For very large datasets, use the generator to process one chart at a time.

# Process charts one at a time (memory efficient)
batch = BatchCalculator.from_natives(sample_natives)

sun_signs = []
for chart in batch.calculate():
    sun = chart.get_object("Sun")
    sun_signs.append(sun.sign)

print(f"Sun signs: {sun_signs}")
Sun signs: ['Capricorn', 'Gemini', 'Sagittarius', 'Pisces', 'Virgo']

2. DataFrame Conversion

Convert charts to pandas DataFrames for analysis. Three schemas are available:

  • charts_to_dataframe: One row per chart (chart-level data)

  • positions_to_dataframe: One row per position (planet positions)

  • aspects_to_dataframe: One row per aspect

2.1 Chart-Level DataFrame

One row per chart with summary data.

# Convert charts to DataFrame
df = charts_to_dataframe(custom_charts)
print(f"DataFrame shape: {df.shape}")
print(f"\nColumns: {list(df.columns)}")
DataFrame shape: (5, 33)

Columns: ['chart_id', 'name', 'datetime_utc', 'julian_day', 'latitude', 'longitude', 'location_name', 'sun_longitude', 'sun_sign', 'sun_sign_degree', 'moon_longitude', 'moon_sign', 'moon_sign_degree', 'moon_phase', 'moon_illumination', 'asc_longitude', 'asc_sign', 'mc_longitude', 'mc_sign', 'fire_count', 'earth_count', 'air_count', 'water_count', 'cardinal_count', 'fixed_count', 'mutable_count', 'sect', 'retrograde_count', 'has_grand_trine', 'has_t_square', 'has_grand_cross', 'has_yod', 'has_stellium']
# View the data
df[
    [
        "name",
        "sun_sign",
        "moon_sign",
        "asc_sign",
        "fire_count",
        "earth_count",
        "air_count",
        "water_count",
    ]
]
name sun_sign moon_sign asc_sign fire_count earth_count air_count water_count
0 Person A Capricorn Scorpio Aries 3 3 3 1
1 Person B Gemini Aries Leo 2 3 3 2
2 Person C Sagittarius Pisces Virgo 2 5 0 3
3 Person D Pisces Gemini Aquarius 2 1 3 4
4 Person E Virgo Taurus Sagittarius 2 5 2 1
# Sun sign distribution
df["sun_sign"].value_counts()
sun_sign
Capricorn      1
Gemini         1
Sagittarius    1
Pisces         1
Virgo          1
Name: count, dtype: int64

2.2 Position-Level DataFrame

One row per celestial position - useful for analyzing specific planets across many charts.

# Convert to positions DataFrame
positions_df = positions_to_dataframe(custom_charts)
print(f"DataFrame shape: {positions_df.shape}")
positions_df.head(10)
DataFrame shape: (100, 13)
chart_id chart_name object_name object_type longitude latitude sign sign_degree house speed is_retrograde declination is_out_of_bounds
0 3ce90358e63e Person A Sun planet 280.581303 0.000226 Capricorn 10.581303 9 1.019454 False -23.015743 False
1 3ce90358e63e Person A Moon planet 225.825072 5.128781 Scorpio 15.825072 7 11.991768 False -11.660509 False
2 3ce90358e63e Person A Mercury planet 272.213612 -1.015072 Capricorn 2.213612 9 1.557362 False -24.434070 True
3 3ce90358e63e Person A Venus planet 241.817697 2.060472 Sagittarius 1.817697 8 1.209279 False -18.504742 False
4 3ce90358e63e Person A Mars planet 328.124902 -1.065186 Aquarius 28.124902 11 0.775680 False -13.124048 False
5 3ce90358e63e Person A Jupiter planet 25.261653 -1.261113 Aries 25.261653 1 0.041465 False 8.598380 False
6 3ce90358e63e Person A Saturn planet 40.391548 -2.443865 Taurus 10.391548 1 -0.019564 True 12.614428 False
7 3ce90358e63e Person A Uranus planet 314.819685 -0.658282 Aquarius 14.819685 11 0.050436 False -17.017199 False
8 3ce90358e63e Person A Neptune planet 303.200427 0.234945 Aquarius 3.200427 10 0.035615 False -19.211575 False
9 3ce90358e63e Person A Pluto planet 251.462095 10.855540 Sagittarius 11.462095 8 0.035101 False -11.394915 False
# Filter to planets only
planets_df = positions_df[positions_df["object_type"] == "planet"]
print(f"Planet positions: {len(planets_df)}")

# Sign distribution for all planets
planets_df["sign"].value_counts()
Planet positions: 50
sign
Capricorn      9
Scorpio        6
Sagittarius    6
Gemini         5
Aquarius       4
Aries          4
Taurus         4
Virgo          4
Pisces         4
Libra          2
Cancer         1
Leo            1
Name: count, dtype: int64
# Retrograde analysis
retrograde_df = planets_df[planets_df["is_retrograde"]]
print(f"Retrograde positions: {len(retrograde_df)}")
retrograde_df[["chart_name", "object_name", "sign", "is_retrograde"]]
Retrograde positions: 10
chart_name object_name sign is_retrograde
6 Person A Saturn Taurus True
25 Person B Jupiter Sagittarius True
27 Person B Uranus Capricorn True
28 Person B Neptune Capricorn True
29 Person B Pluto Scorpio True
45 Person C Jupiter Taurus True
67 Person D Uranus Scorpio True
68 Person D Neptune Sagittarius True
69 Person D Pluto Libra True
86 Person E Saturn Capricorn True

2.3 Aspect-Level DataFrame

One row per aspect - useful for aspect frequency analysis.

# Convert to aspects DataFrame
aspects_df = aspects_to_dataframe(charts_with_aspects)
print(f"DataFrame shape: {aspects_df.shape}")
aspects_df.head(10)
DataFrame shape: (215, 9)
chart_id chart_name object1 object2 aspect_name aspect_degree orb is_applying aspect_type
0 3ce90358e63e Person A Sun Moon Sextile 60.0 5.243769 False longitude
1 3ce90358e63e Person A Sun Saturn Trine 120.0 0.189755 False longitude
2 3ce90358e63e Person A Sun MC Conjunction 0.0 0.136622 None longitude
3 3ce90358e63e Person A Sun Vertex Square 90.0 2.136066 None longitude
4 3ce90358e63e Person A Moon Saturn Opposition 180.0 5.433524 False longitude
5 3ce90358e63e Person A Moon Uranus Square 90.0 1.005387 False longitude
6 3ce90358e63e Person A Moon MC Sextile 60.0 5.107147 None longitude
7 3ce90358e63e Person A Mercury Mars Sextile 60.0 4.088710 False longitude
8 3ce90358e63e Person A Mercury Jupiter Trine 120.0 6.951959 False longitude
9 3ce90358e63e Person A Mercury Vertex Square 90.0 6.231625 None longitude
# Most common aspects
aspects_df["aspect_name"].value_counts()
aspect_name
Square         61
Sextile        53
Trine          53
Opposition     25
Conjunction    23
Name: count, dtype: int64
# Sun-Moon aspects
sun_moon = aspects_df[
    ((aspects_df["object1"] == "Sun") & (aspects_df["object2"] == "Moon"))
    | ((aspects_df["object1"] == "Moon") & (aspects_df["object2"] == "Sun"))
]
print(f"Sun-Moon aspects: {len(sun_moon)}")
sun_moon[["chart_name", "aspect_name", "orb"]]
Sun-Moon aspects: 4
chart_name aspect_name orb
0 Person A Sextile 5.243769
83 Person C Square 0.910136
132 Person D Square 3.660547
173 Person E Trine 5.773051

3. ChartQuery

Filter chart collections by astrological criteria using a fluent API.

Query methods include:

  • where_sun(), where_moon(), where_planet(), where_angle()

  • where_aspect(), where_pattern()

  • where_element_dominant(), where_modality_dominant()

  • where_sect(), where_custom()

3.1 Basic Filtering

# Find charts with Sun in fire signs
fire_sun_charts = (
    ChartQuery(custom_charts).where_sun(sign=["Aries", "Leo", "Sagittarius"]).results()
)
print(f"Charts with Sun in fire signs: {len(fire_sun_charts)}")
for chart in fire_sun_charts:
    print(f"  - {chart.metadata.get('name')}: Sun in {chart.get_object('Sun').sign}")
Charts with Sun in fire signs: 1
  - Person C: Sun in Sagittarius
# Find charts with Mercury retrograde
mercury_rx = (
    ChartQuery(custom_charts).where_planet("Mercury", retrograde=True).results()
)
print(f"Charts with Mercury retrograde: {len(mercury_rx)}")
Charts with Mercury retrograde: 0

3.2 Chained Filters

Combine multiple criteria for complex queries.

# Find day charts with Sun in earth signs
results = (
    ChartQuery(custom_charts)
    .where_sun(sign=["Taurus", "Virgo", "Capricorn"])
    .where_sect("day")
    .results()
)
print(f"Day charts with Sun in earth signs: {len(results)}")
Day charts with Sun in earth signs: 2
# Find charts with Sun-Moon conjunction
conjunctions = (
    ChartQuery(charts_with_aspects)
    .where_aspect("Sun", "Moon", aspect="Conjunction", orb_max=10)
    .results()
)
print(f"Charts with Sun-Moon conjunction: {len(conjunctions)}")
Charts with Sun-Moon conjunction: 0

3.3 Element and Modality Dominance

# Find charts with dominant fire element (4+ planets)
fire_dominant = (
    ChartQuery(custom_charts).where_element_dominant("fire", min_count=4).results()
)
print(f"Charts with fire dominance: {len(fire_dominant)}")

# Find charts with dominant cardinal modality
cardinal_dominant = (
    ChartQuery(custom_charts).where_modality_dominant("cardinal", min_count=4).results()
)
print(f"Charts with cardinal dominance: {len(cardinal_dominant)}")
Charts with fire dominance: 0
Charts with cardinal dominance: 1

3.4 Custom Predicates

Use any custom function to filter charts.

# Find charts with more than 5 aspects
many_aspects = (
    ChartQuery(charts_with_aspects).where_custom(lambda c: len(c.aspects) > 5).results()
)
print(f"Charts with > 5 aspects: {len(many_aspects)}")
Charts with > 5 aspects: 5
# Find charts where Moon is in the same sign as Sun
sun_moon_same = (
    ChartQuery(custom_charts)
    .where_custom(
        lambda c: (
            (s := c.get_object("Sun"))
            and (m := c.get_object("Moon"))
            and s.sign == m.sign
        )
    )
    .results()
)
print(f"Charts with Sun and Moon in same sign: {len(sun_moon_same)}")
for chart in sun_moon_same:
    if sun := chart.get_object("Sun"):
        print(f"  - {chart.metadata.get('name')}: both in {sun.sign}")
Charts with Sun and Moon in same sign: 0

3.5 Result Methods

query = ChartQuery(custom_charts).where_sun(
    sign=["Aries", "Taurus", "Gemini", "Cancer"]
)

# Get count
print(f"Count: {query.count()}")

# Get first result
first = query.first()
if first:
    print(f"First match: {first.metadata.get('name')}")

# Convert to DataFrame
df = query.to_dataframe()
df[["name", "sun_sign", "moon_sign"]]
Count: 1
First match: Person B
name sun_sign moon_sign
0 Person B Gemini Aries

4. ChartStats

Compute aggregate statistics across chart collections.

# Create stats object
stats = ChartStats(custom_charts)
print(f"Analyzing {stats.chart_count} charts")
Analyzing 5 charts

4.1 Element and Modality Distribution

# Element distribution (normalized proportions)
elements = stats.element_distribution()
print("Element Distribution:")
for element, proportion in elements.items():
    print(f"  {element.title()}: {proportion:.1%}")
Element Distribution:
  Fire: 22.0%
  Earth: 34.0%
  Air: 22.0%
  Water: 22.0%
# Modality distribution
modalities = stats.modality_distribution()
print("Modality Distribution:")
for modality, proportion in modalities.items():
    print(f"  {modality.title()}: {proportion:.1%}")
Modality Distribution:
  Cardinal: 32.0%
  Fixed: 30.0%
  Mutable: 38.0%

4.2 Sign Distribution

# Sun sign distribution
sun_signs = stats.sign_distribution("Sun")
print("Sun Sign Distribution:")
for sign, count in sun_signs.items():
    if count > 0:
        print(f"  {sign}: {count}")
Sun Sign Distribution:
  Gemini: 1
  Virgo: 1
  Sagittarius: 1
  Capricorn: 1
  Pisces: 1
# Moon sign distribution
moon_signs = stats.sign_distribution("Moon")
print("Moon Sign Distribution:")
for sign, count in moon_signs.items():
    if count > 0:
        print(f"  {sign}: {count}")
Moon Sign Distribution:
  Aries: 1
  Taurus: 1
  Gemini: 1
  Scorpio: 1
  Pisces: 1

4.3 House Distribution

# Sun house distribution
sun_houses = stats.house_distribution("Sun")
print("Sun House Distribution:")
for house, count in sun_houses.items():
    if count > 0:
        print(f"  House {house}: {count}")
Sun House Distribution:
  House 1: 1
  House 4: 1
  House 9: 2
  House 11: 1

4.4 Aspect Frequency

# Create stats from charts with aspects
stats_aspects = ChartStats(charts_with_aspects)

# Overall aspect frequency
aspect_freq = stats_aspects.aspect_frequency()
print("Aspect Frequency:")
for aspect, count in list(aspect_freq.items())[:5]:
    print(f"  {aspect}: {count}")
Aspect Frequency:
  Square: 61
  Sextile: 53
  Trine: 53
  Opposition: 25
  Conjunction: 23
# Aspect frequency between specific planets
sun_moon_aspects = stats_aspects.aspect_pair_frequency("Sun", "Moon")
print("Sun-Moon Aspect Frequency:")
for aspect, count in sun_moon_aspects.items():
    print(f"  {aspect}: {count}")
Sun-Moon Aspect Frequency:
  Sextile: 1
  Square: 2
  Trine: 1

4.5 Retrograde Frequency

# Retrograde frequency by planet
retro_freq = stats.retrograde_frequency()
print("Retrograde Counts by Planet:")
for planet, count in retro_freq.items():
    print(f"  {planet}: {count}")
Retrograde Counts by Planet:
  Saturn: 2
  Jupiter: 2
  Uranus: 2
  Neptune: 2
  Pluto: 2
# Retrograde rate (normalized)
retro_rate = stats.retrograde_frequency(normalize=True)
print("Retrograde Rate by Planet:")
for planet, rate in retro_rate.items():
    print(f"  {planet}: {rate:.1%}")
Retrograde Rate by Planet:
  Sun: 0.0%
  Moon: 0.0%
  Mercury: 0.0%
  Venus: 0.0%
  Mars: 0.0%
  Jupiter: 40.0%
  Saturn: 40.0%
  Uranus: 40.0%
  Neptune: 40.0%
  Pluto: 40.0%

4.6 Sect Distribution

# Day vs night charts
sect_dist = stats.sect_distribution()
print("Sect Distribution:")
for sect, count in sect_dist.items():
    print(f"  {sect.title()}: {count}")
Sect Distribution:
  Day: 3
  Night: 2

4.7 Cross-Tabulation

Create contingency tables to analyze relationships between variables.

# Sun sign vs Moon sign cross-tabulation
crosstab = stats.cross_tab("sun_sign", "moon_sign")
print("Sun Sign vs Moon Sign:")
crosstab
Sun Sign vs Moon Sign:
moon_sign Aries Gemini Pisces Scorpio Taurus
sun_sign
Capricorn 0 0 0 1 0
Gemini 1 0 0 0 0
Pisces 0 1 0 0 0
Sagittarius 0 0 1 0 0
Virgo 0 0 0 0 1
# Sun sign vs sect
sun_sect = stats.cross_tab("sun_sign", "sect")
print("Sun Sign vs Sect:")
sun_sect
Sun Sign vs Sect:
sect day night
sun_sign
Capricorn 1 0
Gemini 1 0
Pisces 0 1
Sagittarius 0 1
Virgo 1 0

4.8 Summary Report

# Get comprehensive summary
summary = stats.summary()

print(f"Chart Count: {summary['chart_count']}")
print(f"\nElement Distribution: {summary['element_distribution']}")
print(f"\nModality Distribution: {summary['modality_distribution']}")
print(f"\nSect Distribution: {summary['sect_distribution']}")
Chart Count: 5

Element Distribution: {'fire': 0.22, 'earth': 0.34, 'air': 0.22, 'water': 0.22}

Modality Distribution: {'cardinal': 0.32, 'fixed': 0.3, 'mutable': 0.38}

Sect Distribution: {'day': 3, 'night': 2}

5. Export Utilities

Save chart data to files for external analysis.

(Examples in this notebook use tempfile for cookbook purposes; replace with the actual directory you want.)

5.1 CSV Export

import tempfile
from pathlib import Path

# Export chart-level data
with tempfile.TemporaryDirectory() as tmpdir:
    # Charts CSV
    charts_path = Path(tmpdir) / "charts.csv"
    export_csv(custom_charts, charts_path, schema="charts")
    print(f"Exported charts to {charts_path}")

    # Read back and display
    df = pd.read_csv(charts_path)
    print(f"Shape: {df.shape}")
    df.head()
Exported charts to /tmp/tmplp_awbtg/charts.csv
Shape: (5, 33)
# Export positions
with tempfile.TemporaryDirectory() as tmpdir:
    positions_path = Path(tmpdir) / "positions.csv"
    export_csv(custom_charts, positions_path, schema="positions")

    df = pd.read_csv(positions_path)
    print(f"Positions CSV shape: {df.shape}")
    df.head()
Positions CSV shape: (100, 13)

5.2 JSON Export

import json

with tempfile.TemporaryDirectory() as tmpdir:
    # Standard JSON array
    json_path = Path(tmpdir) / "charts.json"
    export_json(custom_charts, json_path)

    with open(json_path) as f:
        data = json.load(f)

    print(f"Exported {len(data)} charts to JSON")
    print(f"Keys in first chart: {list(data[0].keys())[:10]}...")
Exported 5 charts to JSON
Keys in first chart: ['chart_tags', 'datetime', 'location', 'house_systems', 'default_house_system', 'house_placements', 'positions', 'aspects', 'declination_aspects', 'metadata']...
# JSON Lines format (for streaming large datasets)
with tempfile.TemporaryDirectory() as tmpdir:
    jsonl_path = Path(tmpdir) / "charts.jsonl"
    export_json(custom_charts, jsonl_path, lines=True)

    with open(jsonl_path) as f:
        lines = f.readlines()

    print(f"Exported {len(lines)} lines to JSONL")
    print(f"First line preview: {lines[0][:100]}...")
Exported 5 lines to JSONL
First line preview: {"chart_tags": [], "datetime": {"utc": "2000-01-01T17:00:00+00:00", "julian_date": 2451545.208333333...

6. Full Workflow Examples

Complete research workflows combining multiple features.

6.1 Research Question: Element Distribution in Scientists vs Artists

Compare element distributions between different categories.

# Calculate charts for scientists and artists
scientists = BatchCalculator.from_registry(category="scientist").calculate_all()
artists = BatchCalculator.from_registry(category="artist").calculate_all()

print(f"Scientists: {len(scientists)}")
print(f"Artists: {len(artists)}")

# Compare element distributions
if scientists and artists:
    sci_stats = ChartStats(scientists)
    art_stats = ChartStats(artists)

    print("\nElement Distribution Comparison:")
    print(f"{'Element':<12} {'Scientists':<15} {'Artists':<15}")
    print("-" * 42)

    sci_elements = sci_stats.element_distribution()
    art_elements = art_stats.element_distribution()

    for element in ["fire", "earth", "air", "water"]:
        sci_pct = sci_elements.get(element, 0)
        art_pct = art_elements.get(element, 0)
        print(f"{element.title():<12} {sci_pct:>12.1%}   {art_pct:>12.1%}")
Scientists: 20
Artists: 14

Element Distribution Comparison:
Element      Scientists      Artists        
------------------------------------------
Fire                24.0%          27.9%
Earth               23.0%          27.9%
Air                 21.5%          17.1%
Water               31.5%          27.1%

6.2 Research Question: Mercury Retrograde Frequency

Analyze how often Mercury is retrograde in a collection.

# Get all charts
all_charts = BatchCalculator.from_registry().calculate_all()

if all_charts:
    # Query for Mercury retrograde
    mercury_rx_charts = (
        ChartQuery(all_charts).where_planet("Mercury", retrograde=True).results()
    )

    total = len(all_charts)
    rx_count = len(mercury_rx_charts)

    print(f"Total charts: {total}")
    print(f"Mercury Rx charts: {rx_count}")
    print(f"Mercury Rx rate: {rx_count / total:.1%}")
    print("\n(Expected astronomical rate is ~19%)")
Total charts: 211
Mercury Rx charts: 36
Mercury Rx rate: 17.1%

(Expected astronomical rate is ~19%)

6.3 Research Question: Sun-Moon Aspect Distribution

Analyze the distribution of aspects between Sun and Moon.

# Calculate charts with aspects
charts = BatchCalculator.from_registry().with_aspects().calculate_all()

if charts:
    stats = ChartStats(charts)
    sun_moon_aspects = stats.aspect_pair_frequency("Sun", "Moon")

    print("Sun-Moon Aspect Distribution:")
    total_aspects = sum(sun_moon_aspects.values())
    for aspect, count in sorted(sun_moon_aspects.items(), key=lambda x: -x[1]):
        pct = count / total_aspects if total_aspects > 0 else 0
        print(f"  {aspect}: {count} ({pct:.1%})")
Sun-Moon Aspect Distribution:
  Trine: 20 (29.4%)
  Sextile: 19 (27.9%)
  Square: 14 (20.6%)
  Conjunction: 8 (11.8%)
  Opposition: 7 (10.3%)

6.4 Pipeline: Filter, Analyze, Export

A complete data pipeline example.

# 1. Calculate all charts with aspects
charts = (
    BatchCalculator.from_registry()
    .with_aspects()
    .add_analyzer(AspectPatternAnalyzer())
    .calculate_all()
)

print(f"Step 1: Calculated {len(charts)} charts")

# 2. Filter to fire-dominant charts
fire_charts = ChartQuery(charts).where_element_dominant("fire", min_count=4).results()

print(f"Step 2: Found {len(fire_charts)} fire-dominant charts")

# 3. Analyze the subset
if fire_charts:
    fire_stats = ChartStats(fire_charts)
    print("\nStep 3: Analysis of fire-dominant charts:")
    print(f"  Element distribution: {fire_stats.element_distribution()}")
    print(f"  Sect distribution: {fire_stats.sect_distribution()}")

# 4. Convert to DataFrame for further analysis
if fire_charts:
    df = charts_to_dataframe(fire_charts)
    print(f"\nStep 4: DataFrame created with {len(df)} rows")
    display(df[["name", "sun_sign", "fire_count", "sect"]].head())
Step 1: Calculated 211 charts
Step 2: Found 43 fire-dominant charts

Step 3: Analysis of fire-dominant charts:
  Element distribution: {'fire': 0.4372093023255814, 'earth': 0.1511627906976744, 'air': 0.17674418604651163, 'water': 0.23488372093023255}
  Sect distribution: {'day': 34, 'night': 9}

Step 4: DataFrame created with 43 rows
name sun_sign fire_count sect
0 Claude Monet Scorpio 4 day
1 Vincent van Gogh Aries 4 day
2 Wassily Kandinsky Sagittarius 5 night
3 Andy Warhol Leo 7 day
4 Alan Leo Leo 5 day

7. Statistical Analysis Patterns

Stellium prepares your data for analysis. For hypothesis testing and statistical inference, use standard libraries like scipy and statsmodels.

This section demonstrates the handoff pattern: Stellium extracts astrological data -> you apply statistical methods.

# Calculate all charts from the registry for statistical analysis
from scipy import stats as scipy_stats

all_charts = (
    BatchCalculator.from_registry(event_type="birth")
    .with_aspects()
    .add_analyzer(AspectPatternAnalyzer())
    .calculate_all()
)
print(f"Loaded {len(all_charts)} charts from NotableRegistry")
Loaded 190 charts from NotableRegistry

7.1 Chi-Square Test: Sun Sign Distribution

Test whether Sun signs are uniformly distributed (null hypothesis: each sign has equal probability of 1/12).

# Get Sun sign distribution from Stellium
chart_stats = ChartStats(all_charts)
sun_dist = chart_stats.sign_distribution("Sun")

# Observed counts (in zodiac order)
zodiac_order = [
    "Aries",
    "Taurus",
    "Gemini",
    "Cancer",
    "Leo",
    "Virgo",
    "Libra",
    "Scorpio",
    "Sagittarius",
    "Capricorn",
    "Aquarius",
    "Pisces",
]
observed = [sun_dist[sign] for sign in zodiac_order]

# Expected counts under uniform distribution
n_charts = len(all_charts)
expected = [n_charts / 12] * 12

# Chi-square goodness-of-fit test
chi2, p_value = scipy_stats.chisquare(observed, expected)

print("Sun Sign Distribution Test (H₀: uniform distribution)")
print("=" * 55)
print(f"\n{'Sign':<12} {'Observed':>10} {'Expected':>10}")
print("-" * 34)
for sign, obs, exp in zip(zodiac_order, observed, expected, strict=False):
    print(f"{sign:<12} {obs:>10} {exp:>10.1f}")

print(f"\nχ² = {chi2:.3f}")
print(f"p-value = {p_value:.4f}")
print(
    f"\nConclusion: {'Reject H₀ (p < 0.05)' if p_value < 0.05 else 'Cannot reject H₀ (p ≥ 0.05)'}"
)
Sun Sign Distribution Test (H₀: uniform distribution)
=======================================================

Sign           Observed   Expected
----------------------------------
Aries                 8       15.8
Taurus               18       15.8
Gemini               16       15.8
Cancer               18       15.8
Leo                  21       15.8
Virgo                 8       15.8
Libra                17       15.8
Scorpio              17       15.8
Sagittarius          17       15.8
Capricorn            22       15.8
Aquarius             16       15.8
Pisces               12       15.8

χ² = 13.621
p-value = 0.2547

Conclusion: Cannot reject H₀ (p ≥ 0.05)

7.2 Chi-Square Test of Independence: Sun Sign vs Moon Sign

Test whether Sun sign and Moon sign placements are independent of each other.

# Get cross-tabulation from Stellium
crosstab = chart_stats.cross_tab("sun_sign", "moon_sign")

# Reindex to ensure all signs are present
crosstab = crosstab.reindex(index=zodiac_order, columns=zodiac_order, fill_value=0)

print("Sun Sign x Moon Sign Contingency Table:")
display(crosstab)

# Chi-square test of independence
chi2, p_value, dof, expected_freq = scipy_stats.chi2_contingency(crosstab)

print(f"\nχ² = {chi2:.3f}")
print(f"Degrees of freedom = {dof}")
print(f"p-value = {p_value:.4f}")
print(
    f"\nConclusion: {'Reject H₀ - signs may be dependent' if p_value < 0.05 else 'Cannot reject H₀ - signs appear independent'}"
)
Sun Sign x Moon Sign Contingency Table:

χ² = 107.546
Degrees of freedom = 121
p-value = 0.8040

Conclusion: Cannot reject H₀ - signs appear independent
moon_sign Aries Taurus Gemini Cancer Leo Virgo Libra Scorpio Sagittarius Capricorn Aquarius Pisces
sun_sign
Aries 0 0 0 1 0 1 1 1 1 1 1 1
Taurus 2 0 3 2 2 0 2 0 0 3 1 3
Gemini 1 0 1 0 3 3 2 1 1 1 2 1
Cancer 1 2 1 2 1 2 3 2 0 1 2 1
Leo 3 2 5 2 0 0 2 1 3 2 0 1
Virgo 1 0 0 0 0 0 1 2 3 1 0 0
Libra 0 0 1 1 3 3 1 0 2 1 3 2
Scorpio 1 0 1 2 0 1 2 2 3 2 2 1
Sagittarius 2 1 0 0 1 1 3 2 1 3 2 1
Capricorn 1 0 1 4 2 2 3 1 2 1 2 3
Aquarius 1 0 0 2 2 0 1 0 5 2 1 2
Pisces 3 1 1 1 0 0 0 1 3 0 1 1

7.3 Binomial Test: Mercury Retrograde Rate

Test whether Mercury retrograde frequency in our sample differs from the astronomical baseline (~19% of the time).

# Query for Mercury retrograde charts
mercury_rx_charts = (
    ChartQuery(all_charts).where_planet("Mercury", retrograde=True).results()
)

n_total = len(all_charts)
n_rx = len(mercury_rx_charts)
observed_rate = n_rx / n_total
expected_rate = 0.19  # Astronomical baseline

# Binomial test (two-sided)
result = scipy_stats.binomtest(n_rx, n_total, expected_rate, alternative="two-sided")

print("Mercury Retrograde Frequency Test")
print("=" * 45)
print(f"\nTotal charts: {n_total}")
print(f"Mercury Rx charts: {n_rx}")
print(f"Observed rate: {observed_rate:.1%}")
print(f"Expected rate (astronomical): {expected_rate:.1%}")
print(f"\np-value = {result.pvalue:.4f}")
print(f"95% CI: ({result.proportion_ci().low:.1%}, {result.proportion_ci().high:.1%})")
print(
    f"\nConclusion: {'Sample differs from baseline' if result.pvalue < 0.05 else 'Sample consistent with baseline'}"
)
Mercury Retrograde Frequency Test
=============================================

Total charts: 190
Mercury Rx charts: 32
Observed rate: 16.8%
Expected rate (astronomical): 19.0%

p-value = 0.5173
95% CI: (11.8%, 22.9%)

Conclusion: Sample consistent with baseline

7.4 Correlation: Element Counts

Test whether element counts are correlated (e.g., do charts high in fire tend to be low in water?).

# Get chart DataFrame with element counts
df = charts_to_dataframe(all_charts)

# Extract element columns
elements = ["fire_count", "earth_count", "air_count", "water_count"]
element_df = df[elements]

# Correlation matrix
corr_matrix = element_df.corr()

print("Element Count Correlation Matrix:")
print("(Negative correlations expected since planets sum to ~10)")
display(corr_matrix.round(3))

# Test significance of fire-water correlation
fire = df["fire_count"].values
water = df["water_count"].values
r, p_value = scipy_stats.pearsonr(fire, water)

print("\nFire-Water Correlation:")
print(f"  Pearson r = {r:.3f}")
print(f"  p-value = {p_value:.4f}")
print(f"  {'Significant' if p_value < 0.05 else 'Not significant'} at α=0.05")
Element Count Correlation Matrix:
(Negative correlations expected since planets sum to ~10)

Fire-Water Correlation:
  Pearson r = -0.167
  p-value = 0.0215
  Significant at α=0.05
fire_count earth_count air_count water_count
fire_count 1.000 -0.436 -0.265 -0.167
earth_count -0.436 1.000 -0.293 -0.427
air_count -0.265 -0.293 1.000 -0.397
water_count -0.167 -0.427 -0.397 1.000

7.5 Comparing Groups: Scientists vs Artists

Use a Mann-Whitney U test to compare element distributions between categories.

# Calculate charts by category
scientists = BatchCalculator.from_registry(category="scientist").calculate_all()
artists = BatchCalculator.from_registry(category="artist").calculate_all()

# Convert to DataFrames
sci_df = charts_to_dataframe(scientists)
art_df = charts_to_dataframe(artists)

print(f"Scientists: n={len(sci_df)}")
print(f"Artists: n={len(art_df)}")

# Compare water element counts between groups
print("\n" + "=" * 50)
print("Water Element Comparison: Scientists vs Artists")
print("=" * 50)

sci_water = sci_df["water_count"].values
art_water = art_df["water_count"].values

print(
    f"\nScientists - Water count: mean={sci_water.mean():.2f}, median={pd.Series(sci_water).median():.1f}"
)
print(
    f"Artists - Water count: mean={art_water.mean():.2f}, median={pd.Series(art_water).median():.1f}"
)

# Mann-Whitney U test (non-parametric, doesn't assume normal distribution)
stat, p_value = scipy_stats.mannwhitneyu(sci_water, art_water, alternative="two-sided")

print("\nMann-Whitney U test:")
print(f"  U statistic = {stat:.1f}")
print(f"  p-value = {p_value:.4f}")
print(
    f"  {'Groups differ significantly' if p_value < 0.05 else 'No significant difference'}"
)
Scientists: n=20
Artists: n=14

==================================================
Water Element Comparison: Scientists vs Artists
==================================================

Scientists - Water count: mean=3.15, median=3.0
Artists - Water count: mean=2.71, median=3.0

Mann-Whitney U test:
  U statistic = 165.0
  p-value = 0.3797
  No significant difference

7.6 Pattern: Full Statistical Workflow

A complete example combining Stellium data extraction with statistical analysis.

# Research question: Do people with Grand Trines have different element distributions?

# Step 1: Stellium - Query charts with Grand Trines
grand_trine_charts = ChartQuery(all_charts).where_pattern("Grand Trine").results()

no_grand_trine_charts = (
    ChartQuery(all_charts)
    .where_custom(
        lambda c: (
            not any(
                p.name == "Grand Trine" for p in c.metadata.get("aspect_patterns", [])
            )
        )
    )
    .results()
)

print(f"Charts with Grand Trine: {len(grand_trine_charts)}")
print(f"Charts without Grand Trine: {len(no_grand_trine_charts)}")

# Step 2: Stellium - Convert to DataFrames
gt_df = charts_to_dataframe(grand_trine_charts)
no_gt_df = charts_to_dataframe(no_grand_trine_charts)

# Step 3: scipy - Statistical comparison
print("\n" + "=" * 60)
print("Element Distribution: Grand Trine vs No Grand Trine")
print("=" * 60)

for element in ["fire", "earth", "air", "water"]:
    col = f"{element}_count"
    gt_vals = gt_df[col].values
    no_gt_vals = no_gt_df[col].values

    gt_mean = gt_vals.mean()
    no_gt_mean = no_gt_vals.mean()

    stat, p = scipy_stats.mannwhitneyu(gt_vals, no_gt_vals, alternative="two-sided")
    sig = "*" if p < 0.05 else ""

    print(f"\n{element.title()}:")
    print(f"  Grand Trine: mean={gt_mean:.2f}")
    print(f"  No Grand Trine: mean={no_gt_mean:.2f}")
    print(f"  p-value = {p:.4f} {sig}")
Charts with Grand Trine: 52
Charts without Grand Trine: 138

============================================================
Element Distribution: Grand Trine vs No Grand Trine
============================================================

Fire:
  Grand Trine: mean=2.38
  No Grand Trine: mean=2.47
  p-value = 0.3214 

Earth:
  Grand Trine: mean=2.79
  No Grand Trine: mean=2.30
  p-value = 0.1234 

Air:
  Grand Trine: mean=2.08
  No Grand Trine: mean=2.52
  p-value = 0.0250 *

Water:
  Grand Trine: mean=2.75
  No Grand Trine: mean=2.70
  p-value = 0.6876

Summary

The stellium.analysis module provides a complete toolkit for large-scale astrological data analysis:

Component

Purpose

BatchCalculator

Efficiently calculate many charts at once

charts_to_dataframe()

Convert to chart-level DataFrame

positions_to_dataframe()

Convert to position-level DataFrame

aspects_to_dataframe()

Convert to aspect-level DataFrame

ChartQuery

Filter charts by astrological criteria

ChartStats

Compute aggregate statistics

export_csv()

Export to CSV files

export_json()

Export to JSON or JSONL

Statistical analysis is handled by external libraries:

Library

Use Case

scipy.stats

Hypothesis testing (chi-square, t-tests, Mann-Whitney)

statsmodels

Regression, ANOVA, more advanced tests

pingouin

User-friendly statistical tests

Stellium prepares the data; you choose the statistical methodology appropriate for your research question.

For more information, see the Stellium documentation.