Statistical Analysis¶
Calculate hundreds of charts at once, convert them to pandas DataFrames, query them by astrological criteria, and run statistics across the whole set.
The analysis module provides tools for:
Batch chart calculation - Calculate 100s-1000s of charts efficiently
DataFrame conversion - Export charts to pandas DataFrames
Research queries - Filter charts by astrological criteria
Statistical analysis - Aggregate statistics across chart collections
Export utilities - Save to CSV, JSON, Parquet
Installation¶
The analysis module requires pandas (optional dependency):
pip install stellium[analysis]
# Imports
import pandas as pd
from stellium.analysis import (
BatchCalculator,
ChartQuery,
ChartStats,
aspects_to_dataframe,
charts_to_dataframe,
export_csv,
export_json,
positions_to_dataframe,
)
from stellium.core.native import Native
from stellium.engines.patterns import AspectPatternAnalyzer
1. BatchCalculator¶
Efficiently calculate many charts at once. BatchCalculator supports:
Loading from the NotableRegistry (with filters)
Loading from a list of Native objects
Generator-based calculation (memory efficient)
Progress callbacks
1.1 From NotableRegistry¶
Load charts from the built-in database of notable births and events.
# Calculate charts for all notables in the registry
all_charts = BatchCalculator.from_registry().calculate_all()
print(f"Calculated {len(all_charts)} charts")
Calculated 211 charts
# Filter by category
scientist_charts = BatchCalculator.from_registry(category="scientist").calculate_all()
print(f"Calculated {len(scientist_charts)} scientist charts")
Calculated 20 scientist charts
# Multiple filters: verified high-quality data only
verified_charts = BatchCalculator.from_registry(
verified=True, data_quality="AA"
).calculate_all()
print(f"Calculated {len(verified_charts)} verified AA-quality charts")
Calculated 54 verified AA-quality charts
1.2 From Native Objects¶
Calculate charts from your own data.
# Create sample Native objects
sample_natives = [
Native("2000-01-01 12:00", "New York, NY", name="Person A"),
Native("1995-06-21 08:30", "Los Angeles, CA", name="Person B"),
Native("1988-12-15 22:00", "Chicago, IL", name="Person C"),
Native("1975-03-20 06:00", "Seattle, WA", name="Person D"),
Native("1960-09-10 14:30", "Miami, FL", name="Person E"),
]
# Calculate charts
custom_charts = BatchCalculator.from_natives(sample_natives).calculate_all()
print(f"Calculated {len(custom_charts)} custom charts")
Calculated 5 custom charts
1.3 With Aspects and Analyzers¶
Configure the calculation with aspect detection and pattern analysis.
# Calculate with aspects and pattern detection
charts_with_aspects = (
BatchCalculator.from_natives(sample_natives)
.with_aspects() # Enable aspect calculation
.add_analyzer(AspectPatternAnalyzer()) # Detect Grand Trines, T-Squares, etc.
.calculate_all()
)
# Check aspects for first chart
print(f"First chart has {len(charts_with_aspects[0].aspects)} aspects")
First chart has 40 aspects
1.4 Progress Tracking¶
Track progress for long-running batch calculations.
# Define a progress callback
def show_progress(current, total, name):
print(f" [{current}/{total}] Calculating: {name}")
# Calculate with progress tracking
print("Calculating charts with progress:")
charts = (
BatchCalculator.from_natives(sample_natives[:3]) # Just first 3 for demo
.with_progress(show_progress)
.calculate_all()
)
Calculating charts with progress:
[1/3] Calculating: Person A
[2/3] Calculating: Person B
[3/3] Calculating: Person C
1.5 Generator Mode (Memory Efficient)¶
For very large datasets, use the generator to process one chart at a time.
# Process charts one at a time (memory efficient)
batch = BatchCalculator.from_natives(sample_natives)
sun_signs = []
for chart in batch.calculate():
sun = chart.get_object("Sun")
sun_signs.append(sun.sign)
print(f"Sun signs: {sun_signs}")
Sun signs: ['Capricorn', 'Gemini', 'Sagittarius', 'Pisces', 'Virgo']
2. DataFrame Conversion¶
Convert charts to pandas DataFrames for analysis. Three schemas are available:
charts_to_dataframe: One row per chart (chart-level data)
positions_to_dataframe: One row per position (planet positions)
aspects_to_dataframe: One row per aspect
2.1 Chart-Level DataFrame¶
One row per chart with summary data.
# Convert charts to DataFrame
df = charts_to_dataframe(custom_charts)
print(f"DataFrame shape: {df.shape}")
print(f"\nColumns: {list(df.columns)}")
DataFrame shape: (5, 33)
Columns: ['chart_id', 'name', 'datetime_utc', 'julian_day', 'latitude', 'longitude', 'location_name', 'sun_longitude', 'sun_sign', 'sun_sign_degree', 'moon_longitude', 'moon_sign', 'moon_sign_degree', 'moon_phase', 'moon_illumination', 'asc_longitude', 'asc_sign', 'mc_longitude', 'mc_sign', 'fire_count', 'earth_count', 'air_count', 'water_count', 'cardinal_count', 'fixed_count', 'mutable_count', 'sect', 'retrograde_count', 'has_grand_trine', 'has_t_square', 'has_grand_cross', 'has_yod', 'has_stellium']
# View the data
df[
[
"name",
"sun_sign",
"moon_sign",
"asc_sign",
"fire_count",
"earth_count",
"air_count",
"water_count",
]
]
| name | sun_sign | moon_sign | asc_sign | fire_count | earth_count | air_count | water_count | |
|---|---|---|---|---|---|---|---|---|
| 0 | Person A | Capricorn | Scorpio | Aries | 3 | 3 | 3 | 1 |
| 1 | Person B | Gemini | Aries | Leo | 2 | 3 | 3 | 2 |
| 2 | Person C | Sagittarius | Pisces | Virgo | 2 | 5 | 0 | 3 |
| 3 | Person D | Pisces | Gemini | Aquarius | 2 | 1 | 3 | 4 |
| 4 | Person E | Virgo | Taurus | Sagittarius | 2 | 5 | 2 | 1 |
# Sun sign distribution
df["sun_sign"].value_counts()
sun_sign
Capricorn 1
Gemini 1
Sagittarius 1
Pisces 1
Virgo 1
Name: count, dtype: int64
2.2 Position-Level DataFrame¶
One row per celestial position - useful for analyzing specific planets across many charts.
# Convert to positions DataFrame
positions_df = positions_to_dataframe(custom_charts)
print(f"DataFrame shape: {positions_df.shape}")
positions_df.head(10)
DataFrame shape: (100, 13)
| chart_id | chart_name | object_name | object_type | longitude | latitude | sign | sign_degree | house | speed | is_retrograde | declination | is_out_of_bounds | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 3ce90358e63e | Person A | Sun | planet | 280.581303 | 0.000226 | Capricorn | 10.581303 | 9 | 1.019454 | False | -23.015743 | False |
| 1 | 3ce90358e63e | Person A | Moon | planet | 225.825072 | 5.128781 | Scorpio | 15.825072 | 7 | 11.991768 | False | -11.660509 | False |
| 2 | 3ce90358e63e | Person A | Mercury | planet | 272.213612 | -1.015072 | Capricorn | 2.213612 | 9 | 1.557362 | False | -24.434070 | True |
| 3 | 3ce90358e63e | Person A | Venus | planet | 241.817697 | 2.060472 | Sagittarius | 1.817697 | 8 | 1.209279 | False | -18.504742 | False |
| 4 | 3ce90358e63e | Person A | Mars | planet | 328.124902 | -1.065186 | Aquarius | 28.124902 | 11 | 0.775680 | False | -13.124048 | False |
| 5 | 3ce90358e63e | Person A | Jupiter | planet | 25.261653 | -1.261113 | Aries | 25.261653 | 1 | 0.041465 | False | 8.598380 | False |
| 6 | 3ce90358e63e | Person A | Saturn | planet | 40.391548 | -2.443865 | Taurus | 10.391548 | 1 | -0.019564 | True | 12.614428 | False |
| 7 | 3ce90358e63e | Person A | Uranus | planet | 314.819685 | -0.658282 | Aquarius | 14.819685 | 11 | 0.050436 | False | -17.017199 | False |
| 8 | 3ce90358e63e | Person A | Neptune | planet | 303.200427 | 0.234945 | Aquarius | 3.200427 | 10 | 0.035615 | False | -19.211575 | False |
| 9 | 3ce90358e63e | Person A | Pluto | planet | 251.462095 | 10.855540 | Sagittarius | 11.462095 | 8 | 0.035101 | False | -11.394915 | False |
# Filter to planets only
planets_df = positions_df[positions_df["object_type"] == "planet"]
print(f"Planet positions: {len(planets_df)}")
# Sign distribution for all planets
planets_df["sign"].value_counts()
Planet positions: 50
sign
Capricorn 9
Scorpio 6
Sagittarius 6
Gemini 5
Aquarius 4
Aries 4
Taurus 4
Virgo 4
Pisces 4
Libra 2
Cancer 1
Leo 1
Name: count, dtype: int64
# Retrograde analysis
retrograde_df = planets_df[planets_df["is_retrograde"]]
print(f"Retrograde positions: {len(retrograde_df)}")
retrograde_df[["chart_name", "object_name", "sign", "is_retrograde"]]
Retrograde positions: 10
| chart_name | object_name | sign | is_retrograde | |
|---|---|---|---|---|
| 6 | Person A | Saturn | Taurus | True |
| 25 | Person B | Jupiter | Sagittarius | True |
| 27 | Person B | Uranus | Capricorn | True |
| 28 | Person B | Neptune | Capricorn | True |
| 29 | Person B | Pluto | Scorpio | True |
| 45 | Person C | Jupiter | Taurus | True |
| 67 | Person D | Uranus | Scorpio | True |
| 68 | Person D | Neptune | Sagittarius | True |
| 69 | Person D | Pluto | Libra | True |
| 86 | Person E | Saturn | Capricorn | True |
2.3 Aspect-Level DataFrame¶
One row per aspect - useful for aspect frequency analysis.
# Convert to aspects DataFrame
aspects_df = aspects_to_dataframe(charts_with_aspects)
print(f"DataFrame shape: {aspects_df.shape}")
aspects_df.head(10)
DataFrame shape: (215, 9)
| chart_id | chart_name | object1 | object2 | aspect_name | aspect_degree | orb | is_applying | aspect_type | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 3ce90358e63e | Person A | Sun | Moon | Sextile | 60.0 | 5.243769 | False | longitude |
| 1 | 3ce90358e63e | Person A | Sun | Saturn | Trine | 120.0 | 0.189755 | False | longitude |
| 2 | 3ce90358e63e | Person A | Sun | MC | Conjunction | 0.0 | 0.136622 | None | longitude |
| 3 | 3ce90358e63e | Person A | Sun | Vertex | Square | 90.0 | 2.136066 | None | longitude |
| 4 | 3ce90358e63e | Person A | Moon | Saturn | Opposition | 180.0 | 5.433524 | False | longitude |
| 5 | 3ce90358e63e | Person A | Moon | Uranus | Square | 90.0 | 1.005387 | False | longitude |
| 6 | 3ce90358e63e | Person A | Moon | MC | Sextile | 60.0 | 5.107147 | None | longitude |
| 7 | 3ce90358e63e | Person A | Mercury | Mars | Sextile | 60.0 | 4.088710 | False | longitude |
| 8 | 3ce90358e63e | Person A | Mercury | Jupiter | Trine | 120.0 | 6.951959 | False | longitude |
| 9 | 3ce90358e63e | Person A | Mercury | Vertex | Square | 90.0 | 6.231625 | None | longitude |
# Most common aspects
aspects_df["aspect_name"].value_counts()
aspect_name
Square 61
Sextile 53
Trine 53
Opposition 25
Conjunction 23
Name: count, dtype: int64
# Sun-Moon aspects
sun_moon = aspects_df[
((aspects_df["object1"] == "Sun") & (aspects_df["object2"] == "Moon"))
| ((aspects_df["object1"] == "Moon") & (aspects_df["object2"] == "Sun"))
]
print(f"Sun-Moon aspects: {len(sun_moon)}")
sun_moon[["chart_name", "aspect_name", "orb"]]
Sun-Moon aspects: 4
| chart_name | aspect_name | orb | |
|---|---|---|---|
| 0 | Person A | Sextile | 5.243769 |
| 83 | Person C | Square | 0.910136 |
| 132 | Person D | Square | 3.660547 |
| 173 | Person E | Trine | 5.773051 |
3. ChartQuery¶
Filter chart collections by astrological criteria using a fluent API.
Query methods include:
where_sun(),where_moon(),where_planet(),where_angle()where_aspect(),where_pattern()where_element_dominant(),where_modality_dominant()where_sect(),where_custom()
3.1 Basic Filtering¶
# Find charts with Sun in fire signs
fire_sun_charts = (
ChartQuery(custom_charts).where_sun(sign=["Aries", "Leo", "Sagittarius"]).results()
)
print(f"Charts with Sun in fire signs: {len(fire_sun_charts)}")
for chart in fire_sun_charts:
print(f" - {chart.metadata.get('name')}: Sun in {chart.get_object('Sun').sign}")
Charts with Sun in fire signs: 1
- Person C: Sun in Sagittarius
# Find charts with Mercury retrograde
mercury_rx = (
ChartQuery(custom_charts).where_planet("Mercury", retrograde=True).results()
)
print(f"Charts with Mercury retrograde: {len(mercury_rx)}")
Charts with Mercury retrograde: 0
3.2 Chained Filters¶
Combine multiple criteria for complex queries.
# Find day charts with Sun in earth signs
results = (
ChartQuery(custom_charts)
.where_sun(sign=["Taurus", "Virgo", "Capricorn"])
.where_sect("day")
.results()
)
print(f"Day charts with Sun in earth signs: {len(results)}")
Day charts with Sun in earth signs: 2
# Find charts with Sun-Moon conjunction
conjunctions = (
ChartQuery(charts_with_aspects)
.where_aspect("Sun", "Moon", aspect="Conjunction", orb_max=10)
.results()
)
print(f"Charts with Sun-Moon conjunction: {len(conjunctions)}")
Charts with Sun-Moon conjunction: 0
3.3 Element and Modality Dominance¶
# Find charts with dominant fire element (4+ planets)
fire_dominant = (
ChartQuery(custom_charts).where_element_dominant("fire", min_count=4).results()
)
print(f"Charts with fire dominance: {len(fire_dominant)}")
# Find charts with dominant cardinal modality
cardinal_dominant = (
ChartQuery(custom_charts).where_modality_dominant("cardinal", min_count=4).results()
)
print(f"Charts with cardinal dominance: {len(cardinal_dominant)}")
Charts with fire dominance: 0
Charts with cardinal dominance: 1
3.4 Custom Predicates¶
Use any custom function to filter charts.
# Find charts with more than 5 aspects
many_aspects = (
ChartQuery(charts_with_aspects).where_custom(lambda c: len(c.aspects) > 5).results()
)
print(f"Charts with > 5 aspects: {len(many_aspects)}")
Charts with > 5 aspects: 5
# Find charts where Moon is in the same sign as Sun
sun_moon_same = (
ChartQuery(custom_charts)
.where_custom(
lambda c: (
(s := c.get_object("Sun"))
and (m := c.get_object("Moon"))
and s.sign == m.sign
)
)
.results()
)
print(f"Charts with Sun and Moon in same sign: {len(sun_moon_same)}")
for chart in sun_moon_same:
if sun := chart.get_object("Sun"):
print(f" - {chart.metadata.get('name')}: both in {sun.sign}")
Charts with Sun and Moon in same sign: 0
3.5 Result Methods¶
query = ChartQuery(custom_charts).where_sun(
sign=["Aries", "Taurus", "Gemini", "Cancer"]
)
# Get count
print(f"Count: {query.count()}")
# Get first result
first = query.first()
if first:
print(f"First match: {first.metadata.get('name')}")
# Convert to DataFrame
df = query.to_dataframe()
df[["name", "sun_sign", "moon_sign"]]
Count: 1
First match: Person B
| name | sun_sign | moon_sign | |
|---|---|---|---|
| 0 | Person B | Gemini | Aries |
4. ChartStats¶
Compute aggregate statistics across chart collections.
# Create stats object
stats = ChartStats(custom_charts)
print(f"Analyzing {stats.chart_count} charts")
Analyzing 5 charts
4.1 Element and Modality Distribution¶
# Element distribution (normalized proportions)
elements = stats.element_distribution()
print("Element Distribution:")
for element, proportion in elements.items():
print(f" {element.title()}: {proportion:.1%}")
Element Distribution:
Fire: 22.0%
Earth: 34.0%
Air: 22.0%
Water: 22.0%
# Modality distribution
modalities = stats.modality_distribution()
print("Modality Distribution:")
for modality, proportion in modalities.items():
print(f" {modality.title()}: {proportion:.1%}")
Modality Distribution:
Cardinal: 32.0%
Fixed: 30.0%
Mutable: 38.0%
4.2 Sign Distribution¶
# Sun sign distribution
sun_signs = stats.sign_distribution("Sun")
print("Sun Sign Distribution:")
for sign, count in sun_signs.items():
if count > 0:
print(f" {sign}: {count}")
Sun Sign Distribution:
Gemini: 1
Virgo: 1
Sagittarius: 1
Capricorn: 1
Pisces: 1
# Moon sign distribution
moon_signs = stats.sign_distribution("Moon")
print("Moon Sign Distribution:")
for sign, count in moon_signs.items():
if count > 0:
print(f" {sign}: {count}")
Moon Sign Distribution:
Aries: 1
Taurus: 1
Gemini: 1
Scorpio: 1
Pisces: 1
4.3 House Distribution¶
# Sun house distribution
sun_houses = stats.house_distribution("Sun")
print("Sun House Distribution:")
for house, count in sun_houses.items():
if count > 0:
print(f" House {house}: {count}")
Sun House Distribution:
House 1: 1
House 4: 1
House 9: 2
House 11: 1
4.4 Aspect Frequency¶
# Create stats from charts with aspects
stats_aspects = ChartStats(charts_with_aspects)
# Overall aspect frequency
aspect_freq = stats_aspects.aspect_frequency()
print("Aspect Frequency:")
for aspect, count in list(aspect_freq.items())[:5]:
print(f" {aspect}: {count}")
Aspect Frequency:
Square: 61
Sextile: 53
Trine: 53
Opposition: 25
Conjunction: 23
# Aspect frequency between specific planets
sun_moon_aspects = stats_aspects.aspect_pair_frequency("Sun", "Moon")
print("Sun-Moon Aspect Frequency:")
for aspect, count in sun_moon_aspects.items():
print(f" {aspect}: {count}")
Sun-Moon Aspect Frequency:
Sextile: 1
Square: 2
Trine: 1
4.5 Retrograde Frequency¶
# Retrograde frequency by planet
retro_freq = stats.retrograde_frequency()
print("Retrograde Counts by Planet:")
for planet, count in retro_freq.items():
print(f" {planet}: {count}")
Retrograde Counts by Planet:
Saturn: 2
Jupiter: 2
Uranus: 2
Neptune: 2
Pluto: 2
# Retrograde rate (normalized)
retro_rate = stats.retrograde_frequency(normalize=True)
print("Retrograde Rate by Planet:")
for planet, rate in retro_rate.items():
print(f" {planet}: {rate:.1%}")
Retrograde Rate by Planet:
Sun: 0.0%
Moon: 0.0%
Mercury: 0.0%
Venus: 0.0%
Mars: 0.0%
Jupiter: 40.0%
Saturn: 40.0%
Uranus: 40.0%
Neptune: 40.0%
Pluto: 40.0%
4.6 Sect Distribution¶
# Day vs night charts
sect_dist = stats.sect_distribution()
print("Sect Distribution:")
for sect, count in sect_dist.items():
print(f" {sect.title()}: {count}")
Sect Distribution:
Day: 3
Night: 2
4.7 Cross-Tabulation¶
Create contingency tables to analyze relationships between variables.
# Sun sign vs Moon sign cross-tabulation
crosstab = stats.cross_tab("sun_sign", "moon_sign")
print("Sun Sign vs Moon Sign:")
crosstab
Sun Sign vs Moon Sign:
| moon_sign | Aries | Gemini | Pisces | Scorpio | Taurus |
|---|---|---|---|---|---|
| sun_sign | |||||
| Capricorn | 0 | 0 | 0 | 1 | 0 |
| Gemini | 1 | 0 | 0 | 0 | 0 |
| Pisces | 0 | 1 | 0 | 0 | 0 |
| Sagittarius | 0 | 0 | 1 | 0 | 0 |
| Virgo | 0 | 0 | 0 | 0 | 1 |
# Sun sign vs sect
sun_sect = stats.cross_tab("sun_sign", "sect")
print("Sun Sign vs Sect:")
sun_sect
Sun Sign vs Sect:
| sect | day | night |
|---|---|---|
| sun_sign | ||
| Capricorn | 1 | 0 |
| Gemini | 1 | 0 |
| Pisces | 0 | 1 |
| Sagittarius | 0 | 1 |
| Virgo | 1 | 0 |
4.8 Summary Report¶
# Get comprehensive summary
summary = stats.summary()
print(f"Chart Count: {summary['chart_count']}")
print(f"\nElement Distribution: {summary['element_distribution']}")
print(f"\nModality Distribution: {summary['modality_distribution']}")
print(f"\nSect Distribution: {summary['sect_distribution']}")
Chart Count: 5
Element Distribution: {'fire': 0.22, 'earth': 0.34, 'air': 0.22, 'water': 0.22}
Modality Distribution: {'cardinal': 0.32, 'fixed': 0.3, 'mutable': 0.38}
Sect Distribution: {'day': 3, 'night': 2}
5. Export Utilities¶
Save chart data to files for external analysis.
(Examples in this notebook use tempfile for cookbook purposes; replace with the actual directory you want.)
5.1 CSV Export¶
import tempfile
from pathlib import Path
# Export chart-level data
with tempfile.TemporaryDirectory() as tmpdir:
# Charts CSV
charts_path = Path(tmpdir) / "charts.csv"
export_csv(custom_charts, charts_path, schema="charts")
print(f"Exported charts to {charts_path}")
# Read back and display
df = pd.read_csv(charts_path)
print(f"Shape: {df.shape}")
df.head()
Exported charts to /tmp/tmplp_awbtg/charts.csv
Shape: (5, 33)
# Export positions
with tempfile.TemporaryDirectory() as tmpdir:
positions_path = Path(tmpdir) / "positions.csv"
export_csv(custom_charts, positions_path, schema="positions")
df = pd.read_csv(positions_path)
print(f"Positions CSV shape: {df.shape}")
df.head()
Positions CSV shape: (100, 13)
5.2 JSON Export¶
import json
with tempfile.TemporaryDirectory() as tmpdir:
# Standard JSON array
json_path = Path(tmpdir) / "charts.json"
export_json(custom_charts, json_path)
with open(json_path) as f:
data = json.load(f)
print(f"Exported {len(data)} charts to JSON")
print(f"Keys in first chart: {list(data[0].keys())[:10]}...")
Exported 5 charts to JSON
Keys in first chart: ['chart_tags', 'datetime', 'location', 'house_systems', 'default_house_system', 'house_placements', 'positions', 'aspects', 'declination_aspects', 'metadata']...
# JSON Lines format (for streaming large datasets)
with tempfile.TemporaryDirectory() as tmpdir:
jsonl_path = Path(tmpdir) / "charts.jsonl"
export_json(custom_charts, jsonl_path, lines=True)
with open(jsonl_path) as f:
lines = f.readlines()
print(f"Exported {len(lines)} lines to JSONL")
print(f"First line preview: {lines[0][:100]}...")
Exported 5 lines to JSONL
First line preview: {"chart_tags": [], "datetime": {"utc": "2000-01-01T17:00:00+00:00", "julian_date": 2451545.208333333...
6. Full Workflow Examples¶
Complete research workflows combining multiple features.
6.1 Research Question: Element Distribution in Scientists vs Artists¶
Compare element distributions between different categories.
# Calculate charts for scientists and artists
scientists = BatchCalculator.from_registry(category="scientist").calculate_all()
artists = BatchCalculator.from_registry(category="artist").calculate_all()
print(f"Scientists: {len(scientists)}")
print(f"Artists: {len(artists)}")
# Compare element distributions
if scientists and artists:
sci_stats = ChartStats(scientists)
art_stats = ChartStats(artists)
print("\nElement Distribution Comparison:")
print(f"{'Element':<12} {'Scientists':<15} {'Artists':<15}")
print("-" * 42)
sci_elements = sci_stats.element_distribution()
art_elements = art_stats.element_distribution()
for element in ["fire", "earth", "air", "water"]:
sci_pct = sci_elements.get(element, 0)
art_pct = art_elements.get(element, 0)
print(f"{element.title():<12} {sci_pct:>12.1%} {art_pct:>12.1%}")
Scientists: 20
Artists: 14
Element Distribution Comparison:
Element Scientists Artists
------------------------------------------
Fire 24.0% 27.9%
Earth 23.0% 27.9%
Air 21.5% 17.1%
Water 31.5% 27.1%
6.2 Research Question: Mercury Retrograde Frequency¶
Analyze how often Mercury is retrograde in a collection.
# Get all charts
all_charts = BatchCalculator.from_registry().calculate_all()
if all_charts:
# Query for Mercury retrograde
mercury_rx_charts = (
ChartQuery(all_charts).where_planet("Mercury", retrograde=True).results()
)
total = len(all_charts)
rx_count = len(mercury_rx_charts)
print(f"Total charts: {total}")
print(f"Mercury Rx charts: {rx_count}")
print(f"Mercury Rx rate: {rx_count / total:.1%}")
print("\n(Expected astronomical rate is ~19%)")
Total charts: 211
Mercury Rx charts: 36
Mercury Rx rate: 17.1%
(Expected astronomical rate is ~19%)
6.3 Research Question: Sun-Moon Aspect Distribution¶
Analyze the distribution of aspects between Sun and Moon.
# Calculate charts with aspects
charts = BatchCalculator.from_registry().with_aspects().calculate_all()
if charts:
stats = ChartStats(charts)
sun_moon_aspects = stats.aspect_pair_frequency("Sun", "Moon")
print("Sun-Moon Aspect Distribution:")
total_aspects = sum(sun_moon_aspects.values())
for aspect, count in sorted(sun_moon_aspects.items(), key=lambda x: -x[1]):
pct = count / total_aspects if total_aspects > 0 else 0
print(f" {aspect}: {count} ({pct:.1%})")
Sun-Moon Aspect Distribution:
Trine: 20 (29.4%)
Sextile: 19 (27.9%)
Square: 14 (20.6%)
Conjunction: 8 (11.8%)
Opposition: 7 (10.3%)
6.4 Pipeline: Filter, Analyze, Export¶
A complete data pipeline example.
# 1. Calculate all charts with aspects
charts = (
BatchCalculator.from_registry()
.with_aspects()
.add_analyzer(AspectPatternAnalyzer())
.calculate_all()
)
print(f"Step 1: Calculated {len(charts)} charts")
# 2. Filter to fire-dominant charts
fire_charts = ChartQuery(charts).where_element_dominant("fire", min_count=4).results()
print(f"Step 2: Found {len(fire_charts)} fire-dominant charts")
# 3. Analyze the subset
if fire_charts:
fire_stats = ChartStats(fire_charts)
print("\nStep 3: Analysis of fire-dominant charts:")
print(f" Element distribution: {fire_stats.element_distribution()}")
print(f" Sect distribution: {fire_stats.sect_distribution()}")
# 4. Convert to DataFrame for further analysis
if fire_charts:
df = charts_to_dataframe(fire_charts)
print(f"\nStep 4: DataFrame created with {len(df)} rows")
display(df[["name", "sun_sign", "fire_count", "sect"]].head())
Step 1: Calculated 211 charts
Step 2: Found 43 fire-dominant charts
Step 3: Analysis of fire-dominant charts:
Element distribution: {'fire': 0.4372093023255814, 'earth': 0.1511627906976744, 'air': 0.17674418604651163, 'water': 0.23488372093023255}
Sect distribution: {'day': 34, 'night': 9}
Step 4: DataFrame created with 43 rows
| name | sun_sign | fire_count | sect | |
|---|---|---|---|---|
| 0 | Claude Monet | Scorpio | 4 | day |
| 1 | Vincent van Gogh | Aries | 4 | day |
| 2 | Wassily Kandinsky | Sagittarius | 5 | night |
| 3 | Andy Warhol | Leo | 7 | day |
| 4 | Alan Leo | Leo | 5 | day |
7. Statistical Analysis Patterns¶
Stellium prepares your data for analysis. For hypothesis testing and statistical inference, use standard libraries like scipy and statsmodels.
This section demonstrates the handoff pattern: Stellium extracts astrological data -> you apply statistical methods.
# Calculate all charts from the registry for statistical analysis
from scipy import stats as scipy_stats
all_charts = (
BatchCalculator.from_registry(event_type="birth")
.with_aspects()
.add_analyzer(AspectPatternAnalyzer())
.calculate_all()
)
print(f"Loaded {len(all_charts)} charts from NotableRegistry")
Loaded 190 charts from NotableRegistry
7.1 Chi-Square Test: Sun Sign Distribution¶
Test whether Sun signs are uniformly distributed (null hypothesis: each sign has equal probability of 1/12).
# Get Sun sign distribution from Stellium
chart_stats = ChartStats(all_charts)
sun_dist = chart_stats.sign_distribution("Sun")
# Observed counts (in zodiac order)
zodiac_order = [
"Aries",
"Taurus",
"Gemini",
"Cancer",
"Leo",
"Virgo",
"Libra",
"Scorpio",
"Sagittarius",
"Capricorn",
"Aquarius",
"Pisces",
]
observed = [sun_dist[sign] for sign in zodiac_order]
# Expected counts under uniform distribution
n_charts = len(all_charts)
expected = [n_charts / 12] * 12
# Chi-square goodness-of-fit test
chi2, p_value = scipy_stats.chisquare(observed, expected)
print("Sun Sign Distribution Test (H₀: uniform distribution)")
print("=" * 55)
print(f"\n{'Sign':<12} {'Observed':>10} {'Expected':>10}")
print("-" * 34)
for sign, obs, exp in zip(zodiac_order, observed, expected, strict=False):
print(f"{sign:<12} {obs:>10} {exp:>10.1f}")
print(f"\nχ² = {chi2:.3f}")
print(f"p-value = {p_value:.4f}")
print(
f"\nConclusion: {'Reject H₀ (p < 0.05)' if p_value < 0.05 else 'Cannot reject H₀ (p ≥ 0.05)'}"
)
Sun Sign Distribution Test (H₀: uniform distribution)
=======================================================
Sign Observed Expected
----------------------------------
Aries 8 15.8
Taurus 18 15.8
Gemini 16 15.8
Cancer 18 15.8
Leo 21 15.8
Virgo 8 15.8
Libra 17 15.8
Scorpio 17 15.8
Sagittarius 17 15.8
Capricorn 22 15.8
Aquarius 16 15.8
Pisces 12 15.8
χ² = 13.621
p-value = 0.2547
Conclusion: Cannot reject H₀ (p ≥ 0.05)
7.2 Chi-Square Test of Independence: Sun Sign vs Moon Sign¶
Test whether Sun sign and Moon sign placements are independent of each other.
# Get cross-tabulation from Stellium
crosstab = chart_stats.cross_tab("sun_sign", "moon_sign")
# Reindex to ensure all signs are present
crosstab = crosstab.reindex(index=zodiac_order, columns=zodiac_order, fill_value=0)
print("Sun Sign x Moon Sign Contingency Table:")
display(crosstab)
# Chi-square test of independence
chi2, p_value, dof, expected_freq = scipy_stats.chi2_contingency(crosstab)
print(f"\nχ² = {chi2:.3f}")
print(f"Degrees of freedom = {dof}")
print(f"p-value = {p_value:.4f}")
print(
f"\nConclusion: {'Reject H₀ - signs may be dependent' if p_value < 0.05 else 'Cannot reject H₀ - signs appear independent'}"
)
Sun Sign x Moon Sign Contingency Table:
χ² = 107.546
Degrees of freedom = 121
p-value = 0.8040
Conclusion: Cannot reject H₀ - signs appear independent
| moon_sign | Aries | Taurus | Gemini | Cancer | Leo | Virgo | Libra | Scorpio | Sagittarius | Capricorn | Aquarius | Pisces |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| sun_sign | ||||||||||||
| Aries | 0 | 0 | 0 | 1 | 0 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| Taurus | 2 | 0 | 3 | 2 | 2 | 0 | 2 | 0 | 0 | 3 | 1 | 3 |
| Gemini | 1 | 0 | 1 | 0 | 3 | 3 | 2 | 1 | 1 | 1 | 2 | 1 |
| Cancer | 1 | 2 | 1 | 2 | 1 | 2 | 3 | 2 | 0 | 1 | 2 | 1 |
| Leo | 3 | 2 | 5 | 2 | 0 | 0 | 2 | 1 | 3 | 2 | 0 | 1 |
| Virgo | 1 | 0 | 0 | 0 | 0 | 0 | 1 | 2 | 3 | 1 | 0 | 0 |
| Libra | 0 | 0 | 1 | 1 | 3 | 3 | 1 | 0 | 2 | 1 | 3 | 2 |
| Scorpio | 1 | 0 | 1 | 2 | 0 | 1 | 2 | 2 | 3 | 2 | 2 | 1 |
| Sagittarius | 2 | 1 | 0 | 0 | 1 | 1 | 3 | 2 | 1 | 3 | 2 | 1 |
| Capricorn | 1 | 0 | 1 | 4 | 2 | 2 | 3 | 1 | 2 | 1 | 2 | 3 |
| Aquarius | 1 | 0 | 0 | 2 | 2 | 0 | 1 | 0 | 5 | 2 | 1 | 2 |
| Pisces | 3 | 1 | 1 | 1 | 0 | 0 | 0 | 1 | 3 | 0 | 1 | 1 |
7.3 Binomial Test: Mercury Retrograde Rate¶
Test whether Mercury retrograde frequency in our sample differs from the astronomical baseline (~19% of the time).
# Query for Mercury retrograde charts
mercury_rx_charts = (
ChartQuery(all_charts).where_planet("Mercury", retrograde=True).results()
)
n_total = len(all_charts)
n_rx = len(mercury_rx_charts)
observed_rate = n_rx / n_total
expected_rate = 0.19 # Astronomical baseline
# Binomial test (two-sided)
result = scipy_stats.binomtest(n_rx, n_total, expected_rate, alternative="two-sided")
print("Mercury Retrograde Frequency Test")
print("=" * 45)
print(f"\nTotal charts: {n_total}")
print(f"Mercury Rx charts: {n_rx}")
print(f"Observed rate: {observed_rate:.1%}")
print(f"Expected rate (astronomical): {expected_rate:.1%}")
print(f"\np-value = {result.pvalue:.4f}")
print(f"95% CI: ({result.proportion_ci().low:.1%}, {result.proportion_ci().high:.1%})")
print(
f"\nConclusion: {'Sample differs from baseline' if result.pvalue < 0.05 else 'Sample consistent with baseline'}"
)
Mercury Retrograde Frequency Test
=============================================
Total charts: 190
Mercury Rx charts: 32
Observed rate: 16.8%
Expected rate (astronomical): 19.0%
p-value = 0.5173
95% CI: (11.8%, 22.9%)
Conclusion: Sample consistent with baseline
7.4 Correlation: Element Counts¶
Test whether element counts are correlated (e.g., do charts high in fire tend to be low in water?).
# Get chart DataFrame with element counts
df = charts_to_dataframe(all_charts)
# Extract element columns
elements = ["fire_count", "earth_count", "air_count", "water_count"]
element_df = df[elements]
# Correlation matrix
corr_matrix = element_df.corr()
print("Element Count Correlation Matrix:")
print("(Negative correlations expected since planets sum to ~10)")
display(corr_matrix.round(3))
# Test significance of fire-water correlation
fire = df["fire_count"].values
water = df["water_count"].values
r, p_value = scipy_stats.pearsonr(fire, water)
print("\nFire-Water Correlation:")
print(f" Pearson r = {r:.3f}")
print(f" p-value = {p_value:.4f}")
print(f" {'Significant' if p_value < 0.05 else 'Not significant'} at α=0.05")
Element Count Correlation Matrix:
(Negative correlations expected since planets sum to ~10)
Fire-Water Correlation:
Pearson r = -0.167
p-value = 0.0215
Significant at α=0.05
| fire_count | earth_count | air_count | water_count | |
|---|---|---|---|---|
| fire_count | 1.000 | -0.436 | -0.265 | -0.167 |
| earth_count | -0.436 | 1.000 | -0.293 | -0.427 |
| air_count | -0.265 | -0.293 | 1.000 | -0.397 |
| water_count | -0.167 | -0.427 | -0.397 | 1.000 |
7.5 Comparing Groups: Scientists vs Artists¶
Use a Mann-Whitney U test to compare element distributions between categories.
# Calculate charts by category
scientists = BatchCalculator.from_registry(category="scientist").calculate_all()
artists = BatchCalculator.from_registry(category="artist").calculate_all()
# Convert to DataFrames
sci_df = charts_to_dataframe(scientists)
art_df = charts_to_dataframe(artists)
print(f"Scientists: n={len(sci_df)}")
print(f"Artists: n={len(art_df)}")
# Compare water element counts between groups
print("\n" + "=" * 50)
print("Water Element Comparison: Scientists vs Artists")
print("=" * 50)
sci_water = sci_df["water_count"].values
art_water = art_df["water_count"].values
print(
f"\nScientists - Water count: mean={sci_water.mean():.2f}, median={pd.Series(sci_water).median():.1f}"
)
print(
f"Artists - Water count: mean={art_water.mean():.2f}, median={pd.Series(art_water).median():.1f}"
)
# Mann-Whitney U test (non-parametric, doesn't assume normal distribution)
stat, p_value = scipy_stats.mannwhitneyu(sci_water, art_water, alternative="two-sided")
print("\nMann-Whitney U test:")
print(f" U statistic = {stat:.1f}")
print(f" p-value = {p_value:.4f}")
print(
f" {'Groups differ significantly' if p_value < 0.05 else 'No significant difference'}"
)
Scientists: n=20
Artists: n=14
==================================================
Water Element Comparison: Scientists vs Artists
==================================================
Scientists - Water count: mean=3.15, median=3.0
Artists - Water count: mean=2.71, median=3.0
Mann-Whitney U test:
U statistic = 165.0
p-value = 0.3797
No significant difference
7.6 Pattern: Full Statistical Workflow¶
A complete example combining Stellium data extraction with statistical analysis.
# Research question: Do people with Grand Trines have different element distributions?
# Step 1: Stellium - Query charts with Grand Trines
grand_trine_charts = ChartQuery(all_charts).where_pattern("Grand Trine").results()
no_grand_trine_charts = (
ChartQuery(all_charts)
.where_custom(
lambda c: (
not any(
p.name == "Grand Trine" for p in c.metadata.get("aspect_patterns", [])
)
)
)
.results()
)
print(f"Charts with Grand Trine: {len(grand_trine_charts)}")
print(f"Charts without Grand Trine: {len(no_grand_trine_charts)}")
# Step 2: Stellium - Convert to DataFrames
gt_df = charts_to_dataframe(grand_trine_charts)
no_gt_df = charts_to_dataframe(no_grand_trine_charts)
# Step 3: scipy - Statistical comparison
print("\n" + "=" * 60)
print("Element Distribution: Grand Trine vs No Grand Trine")
print("=" * 60)
for element in ["fire", "earth", "air", "water"]:
col = f"{element}_count"
gt_vals = gt_df[col].values
no_gt_vals = no_gt_df[col].values
gt_mean = gt_vals.mean()
no_gt_mean = no_gt_vals.mean()
stat, p = scipy_stats.mannwhitneyu(gt_vals, no_gt_vals, alternative="two-sided")
sig = "*" if p < 0.05 else ""
print(f"\n{element.title()}:")
print(f" Grand Trine: mean={gt_mean:.2f}")
print(f" No Grand Trine: mean={no_gt_mean:.2f}")
print(f" p-value = {p:.4f} {sig}")
Charts with Grand Trine: 52
Charts without Grand Trine: 138
============================================================
Element Distribution: Grand Trine vs No Grand Trine
============================================================
Fire:
Grand Trine: mean=2.38
No Grand Trine: mean=2.47
p-value = 0.3214
Earth:
Grand Trine: mean=2.79
No Grand Trine: mean=2.30
p-value = 0.1234
Air:
Grand Trine: mean=2.08
No Grand Trine: mean=2.52
p-value = 0.0250 *
Water:
Grand Trine: mean=2.75
No Grand Trine: mean=2.70
p-value = 0.6876
Summary¶
The stellium.analysis module provides a complete toolkit for large-scale astrological data analysis:
Component |
Purpose |
|---|---|
|
Efficiently calculate many charts at once |
|
Convert to chart-level DataFrame |
|
Convert to position-level DataFrame |
|
Convert to aspect-level DataFrame |
|
Filter charts by astrological criteria |
|
Compute aggregate statistics |
|
Export to CSV files |
|
Export to JSON or JSONL |
Statistical analysis is handled by external libraries:
Library |
Use Case |
|---|---|
|
Hypothesis testing (chi-square, t-tests, Mann-Whitney) |
|
Regression, ANOVA, more advanced tests |
|
User-friendly statistical tests |
Stellium prepares the data; you choose the statistical methodology appropriate for your research question.
For more information, see the Stellium documentation.