Track Record Reset: Clean Slate, Daily Rates, and Clearer Money Weather
The old track record answered a useful but narrow question: did the daily rate call point the right way?
That still matters. But it was never enough.
So we reset the public scorecard.
The current-era BlueSkyFI rate signal track record starts on May 19, 2026, after we moved to cleaner daily mortgage-rate data. The older rows are not deleted. They are preserved as a pre-reset archive for accountability, but they no longer count in the current accuracy numbers.
We also paused the HeyGen avatar workflow to keep operating costs under control. That means the public scorecard is no longer framed as an avatar's personal track record. It is now the BlueSkyFI rate signal track record: cleaner name, cleaner source, cleaner sample.
BlueSkyFI is now built around a broader Money Weather question: what should a household do today with cash, debt, housing, job risk, and financial-independence progress in mind?
So the track record had to evolve. It now measures rate-call calibration, the evidence behind the signal, and the new Money Weather factors that turn market news into subscriber actions.
Updated July 31, 2026: the admin dashboard now tracks the human-readable call trail, recent versus prior accuracy, rolling and cumulative scores, recommendation-level accuracy, Gemini probability forecasts, and fresh public-expert agreement. The goal is not just to ask "was Gemini involved?" It is to show what Gemini recommended, whether hard-data guardrails changed it, and whether the final result improved.
The short version
We moved from a narrow mortgage-rate scoreboard to a fuller Money Weather measurement system, then reset the public scorecard once the daily mortgage-rate source went live. The forecast now reads real-world signals, turns them into everyday watchpoints, tracks confidence calibration, and keeps the pre-reset rows archived outside the current metrics.
The August measurement contract
The daily rate call is now a five-business-day forecast with three explicit layers:
- Baseline: the recommendation produced from BlueSkyFI's deterministic market, rate, spread, event-risk, and borrower-margin rules.
- Gemini: a probability forecast for rates up, flat, or down, an expected basis-point move, and a matching lock, hold, or float recommendation.
- Final: the published recommendation after freshness, hard-data alignment, event-risk, track-record, and confidence guardrails are applied.
The Gemini packet cannot influence the final action unless its structured data read supports the same direction. Fresh public expert commentary can corroborate the call, but it cannot replace the data read.
For expert benchmarking, BlueSkyFI accepts no more than three public, dated reads from the prior 72 hours. That can include a clearly public Barry Habib or MBS Highway stance, Mortgage News Daily, HousingWire, Reuters, or a primary Fed or Treasury source. Old videos, annual forecasts, marketing pages, and subscriber-only commentary do not become today's lock/float call.
The dashboard records:
- the rates-up, flat, and down probabilities,
- expected move in basis points,
- baseline, Gemini, and final actions,
- whether Gemini changed the baseline,
- whether the final action matched Gemini,
- whether a fresh expert benchmark existed,
- whether Gemini and the expert consensus agreed,
- and whether a fresh public Barry Habib or MBS Highway read matched.
Agreement is not accuracy. The five-business-day rate outcome still grades the call. Expert agreement is useful because it tells us whether Gemini found and interpreted the same public market evidence as a credible rate desk, not because any expert gets to mark the homework.
What changed
Five upgrades are live.
1) The public scorecard starts clean
The current public track record now starts on May 19, 2026. That is the reset line.
Why reset instead of blending everything together?
- The older rows were produced before the cleaner daily rate data was in place.
- The old sample was useful as a debugging trail, but mixing it with the new source would make the current score harder to read.
- Keeping the archive visible is more honest than deleting weak history or pretending the methodology never changed.
So the dashboard now separates:
- Current-era metrics: rows from May 19, 2026 forward.
- Pre-reset archive: rows through May 18, 2026, preserved for audit history.
That is the balance we want: accountability without dragging a changed methodology through the new sample.
2) The scorecard is more precise
The track record now separates different questions that should never have been blended together:
- Directional accuracy: when we make a lock/float call, did the market move in that direction?
- All-call accuracy: how did every graded call perform, including holds?
- Confident-call accuracy: when confidence is 70% or higher, does the result justify that confidence?
- Brier score: are our confidence numbers calibrated, or just decoration?
- Bad-rate filtering: obviously wrong rate inputs are excluded instead of quietly distorting the read.
This does not make grading easier. It makes grading more useful.
If the system says "high confidence," the dashboard should eventually prove whether that confidence means something. If it does not, we tune it.
3) The signal set is wider
The old setup leaned too heavily on the data that was easiest to fetch. That made it clean, but incomplete.
The updated system starts with the subscriber's actual day, then layers market context around it:
- Cash and debt pressure: high-yield savings yields, short Treasury yields, credit-card APR, CPI/PCE inflation, savings rate, debt drag, and emergency-fund context.
- Income and labor pressure: unemployment, jobless-claims trend, wage pressure, volatility, credit spreads, and market stress.
- FI and portfolio context: FI progress, spending flexibility, savings-rate momentum, and whether market volatility argues for patience or action.
- Housing and rent pressure: existing-home months supply, builder price-cut share, home-price pressure, rent pressure, wage pressure, affordability momentum, and profile-level housing exposure.
- Mortgage-rate context: mortgage rates, Treasury yields, mortgage spreads, yield-curve shape, applications, FOMC tone, Treasury auction demand, term premium, Treasury issuance pressure, foreign demand, and shutdown/debt-ceiling risk.
That gives the forecast more texture. A hot CPI print, a weaker labor trend, improving inventory, and a widening mortgage spread should not all collapse into one generic "rates" story. They affect different households differently.
4) The daily report now remembers what changed
New daily reports preserve more than the final recommendation. They remember:
- The subscriber's Money Weather score and label.
- The daily anthem and recommendation set.
- The market signals that influenced the recommendation.
- The profile details that shaped the action plan.
- Metric packs such as Housing Stress Index and Affordability Velocity.
That matters because the best product is not a single forecast. It is a history of what changed, why the recommendation changed, and whether the next step still fits the user's life.
5) The admin scorecard now shows the human trail
Aggregate accuracy is useful, but it can hide the lived story of the product.
The admin dashboard now keeps a daily call ledger that shows:
- The date of the call.
- The lock, float, or hold recommendation.
- The confidence level.
- The observed mortgage-rate move once the call can be graded.
- Whether the call was right, wrong, or still pending.
- The human rationale that was shown at call time.
- The active Gemini evidence lanes, such as Fed and policy, rates and bonds, mortgage and housing, labor/inflation proxies, or risk and credit.
It also compares the most recent graded calls against the previous batch and charts rolling versus cumulative accuracy. That makes the improvement question visible: are new guardrails, better evidence, and Gemini lane gating actually improving the hit rate, or just making the system feel smarter?
This is the main standard now: keep the daily trail readable, measure the score over time, and tune down any evidence lane that is not earning its keep.
The new Money Weather factors
The old public track record was mostly about mortgage timing. The new system still respects housing timing, but the main product now starts with the factors that affect a subscriber's actual day.
Cash, debt, and income pressure
A subscriber does not live inside a mortgage chart. The daily report now gives more weight to:
- Credit-card APR and debt drag.
- High-yield savings and short Treasury yields.
- CPI/PCE inflation pressure.
- Unemployment and jobless-claims trend.
- Cash runway and emergency-fund profile.
- Savings rate, FI progress, and debt-to-income pressure.
This is the product direction: rate timing is one input, not the center of gravity.
Housing Stress Index
Housing Stress Index is a 0-100 stress score. Higher means more strain on affordability, reserves, and timing decisions.
It combines:
- Mortgage-rate pressure.
- Mortgage spread pressure.
- Yield-curve pressure.
- Labor-market pressure.
- Volatility pressure.
This helps answer: is the housing market merely expensive, or is it actively pressuring the next move?
Affordability Velocity
Affordability Velocity measures whether affordability is improving or worsening fast enough to change behavior.
It combines:
- Housing affordability level.
- Monthly affordability momentum.
- Home-price pressure.
- Rent versus wage pressure.
- Mortgage-rate pressure.
- Profile context when the user is signed in.
This helps answer: should the user protect margin, shrink the target, or move while inventory gives them leverage?
The signal has to match the decision
BlueSkyFI should not ask you to trust hidden machinery. The daily report looks at public market signals, checks whether they are fresh enough to matter, and turns them into plain-language guidance for cash, debt, housing, and financial-independence decisions.
The signal starts with sources people can recognize: inflation and income data, labor releases, Treasury yields, mortgage-rate benchmarks, housing inventory, volatility, cash yields, and debt-cost readings. The point is not to bury you in charts. The point is to answer the question a user actually brought to the report.
If the decision is cash, the report should focus on yield, inflation, job risk, and runway. If the decision is debt, the report should focus on APR pressure and payoff tradeoffs. If the decision is housing, the report should focus on payment, inventory, affordability, and timing risk. If the decision is financial independence, the report should focus on savings rate, market volatility, inflation, and timeline resilience.
The important part for users is simple:
- You see what changed in the market.
- You see why it matters for your household.
- You can compare today's recommendation with earlier recommendations.
- The track record keeps the rate-signal side accountable, while report history keeps the household decision trail visible.
That makes the report more useful over time without making you learn how the machinery works.
How we look ahead without pretending
BlueSkyFI is trying to be early, not dramatic.
That means the report can look beyond today's headline, but it has to stay honest. A useful forecast has three parts:
- Observed data: what actually changed in the latest public data.
- Everyday scenario: what that change could mean if it keeps going.
- Trigger point: what next reading would confirm or cancel the scenario.
For example, sticky inflation does not automatically mean "rates will stay high forever." It means cash budgets, debt payoff, and housing payments need more margin until inflation and rates actually cool. A softer labor market does not automatically mean "good news." It may help rates, but it can also raise job-risk pressure.
That is the 2027 lens. We are not trying to call one perfect future. We are watching the signals that can shape the next 6 to 18 months:
- Inflation and wage growth: is purchasing power improving or still leaking?
- Cash yields and credit-card APRs: is the safe return still useful, and is debt still punitive?
- Treasury yields and mortgage rates: are borrowing costs easing enough to change housing or refi math?
- Jobless claims and unemployment: is rate relief coming with more income risk?
- Housing inventory and affordability: is the market giving buyers real room, or just better headlines?
- Volatility and market breadth: is the portfolio environment calm because risk is low, or calm because investors are ignoring slow pressure?
In ELI5 terms: we check the weather, then we ask what would make you carry an umbrella, leave earlier, or stay home.
What will keep improving
This system is not frozen. It is meant to be tuned.
Here is how the report keeps getting sharper:
- Watch the current market backdrop.
- Translate it into household pressure: cash, debt, job, housing, market risk, and FI path.
- Save the recommendation so you can see what changed.
- Grade the rate-signal side after enough time passes.
- Track which watchpoints kept mattering and which ones were noise.
- Keep the parts that help and tone down the parts that do not.
That is the right posture. Not "trust the score alone." Not "trust the vibe." Trust the scorecard as it accumulates enough clean observations.
How to read the dashboard
Use the dashboard in layers:
- Directional accuracy tells you whether lock/float calls are moving the right way.
- Confident-call accuracy tells you whether high-conviction calls deserve more attention.
- Brier score tells you whether confidence is calibrated.
- Daily call ledger shows the human record: what we said, why we said it, what rates did afterward, and whether the call graded right.
- Rolling score trend compares recent calls with prior calls so improvement or deterioration is visible.
- Gemini lane lift tells you whether specific evidence families are adding value instead of treating Gemini as one giant black box.
- Metric history shows how Housing Stress and Affordability Velocity evolve over time.
Early windows will still show "pending" when the sample is too small. That is a feature. Thin data should not cosplay as certainty.
What this means for users
For a subscriber, the practical upgrade is simple:
- The report is less likely to over-index on mortgage rates.
- The action plan sees cash, debt, income, housing, and FI context together.
- The recommendation history becomes useful because it keeps the anthem, market context, and next steps together.
- BlueSkyFI can track whether the new metrics are helping.
If you are checking BlueSkyFI before a housing, debt, cash, or FI decision, this is the direction you want: more evidence, clearer calibration, fewer one-number stories.
Where to look
- Public track record - current-era directional, all-call, confident, Brier, calibration buckets, and pre-reset archive summary.
- Methodology - grading rules in detail.
- BlueSky Report - subscriber Money Weather and daily recommendations.
- Report history - saved daily reports once subscriber history is enabled.
- August FI playbook - the current household-market interpretation and action triggers.
- August MBS playbook - the current five-business-day rate setup, payment exposure, and desk triggers.
The point is not to pretend every daily call will be right. The point is to make every update measurable, explainable, and more useful than the last one.
