Overview01
- 01
Diagnosed the root cause of a stalled inverse-FEA predictor by benchmarking 7 ML models (Random Forest, XGBoost, HGB/ExtraTrees ensembles, PCA pipelines) and running paired clean-vs-noisy signal analysis. This proved the R² ≈ 0.10 ceiling was a data limitation rather than a model limitation and redirected the team away from futile model tuning.
- 02
Raised earthquake-detection accuracy to 0.992 ROC-AUC (vs. 0.550 baseline) and magnitude-estimation R² to 0.840 by training PyTorch CNN→Transformer and GNN ensembles on 5,800+ multi-station SCEDC waveform windows.
- 03
Built a real-time daemon streaming 10 SeedLink stations that runs continuous CNN→Transformer detection plus GNN-fused multi-station magnitude estimation. It confirms events with graded multi-station coincidence and geographic move-out checks before dispatching FCM push alerts, and it's deployed on an Oracle Cloud VM under systemd.
Diagnosing the inverse-FEA ceiling02
An inverse finite-element problem runs simulation backwards: given a structure's measured response, recover the material properties that produced it. One target in the model behind the team's IEEE research paper was stuck at R² ≈ 0.10, and the instinct was to keep tuning models. I set out to find out whether tuning could help at all.
This work is part of an unpublished paper, so its figures and detailed results are held back until publication. The method is below.
Step 1 · Benchmark seven models
If the bottleneck is the model, different inductive biases should give different answers. I benchmarked seven model configurations under the same cross-validation: a baseline random forest, chained random forests, and my own models, including boosted / ExtraTrees ensembles and PCA pipelines. Every one landed on the same ceiling for the hard target, while the easier targets were predicted well by all of them.
Seven very different models, one ceiling. That's the data talking.
Step 2 · Paired clean vs. noisy data
The decisive test was to train on clean data and on noisy data side by side. If a target is recoverable from clean data but not once realistic noise is added, then no model trained on noisy data will get it back, however it's tuned. That's exactly what happened.
Step 3 · Where the noise lives
Next I measured how much noise each group of measurements carries relative to its signal, and ran an ablation using only the cleanest inputs. Dropping the noisiest inputs didn't rescue the target.
Step 4 · How fragile is each target?
Finally I swept the amount of measurement noise and tracked R² for each target: an identifiability curve showing how much noise each target can tolerate. At the real measurement noise, the hard target can't be recovered.
Could an eighth model have fixed it?
No. On clean data the target is recoverable. With the real measurement noise, no model gets it back.
The outcome
Every line of evidence pointed the same way: the R² ≈ 0.10 ceiling was a data limitation, not a model limitation. That redirected the team from model tuning to the inputs: what's measured, and at what signal quality. A negative result like this saves weeks.
Building SeismicSoCal03
The second half of the internship became SeismicSoCal, deep-learning seismology for Southern California, now deployed live. The full case study has the details. In short:
How I approach research04
Both halves of this internship (and my other research) follow the same loop. Here is each step, with what it looked like in practice.
Research is a loop, not a line.
- 01
Frame the question
Turn a vague problem into one that evidence can answer.
- CBU: is the R² ≈ 0.10 ceiling a model problem or a data problem?
- SeismicSoCal: can a learned detector tell real quakes from noise better than the classic STA/LTA trigger?
- 02
Know the field & pick baselines
Start from what practitioners already use, so a result means something.
- STA/LTA for detection, amplitude + distance for magnitude, a GMPE-style estimate for shaking.
- Persistence for forecasting (Berkeley Lab), a baseline random forest for inverse-FEA.
- 03
Design the experiment
Decide how results will be judged before running anything.
- Chronological 70/15/15 splits and walk-forward folds, so no model ever sees the future.
- Thresholds tuned on validation only; a paired clean-vs-noisy design that isolates noise as the one variable.
- 04
Collect & check the data
Most bad results start as bad data.
- 5,800+ SCEDC waveform windows labelled against the USGS catalog.
- A quality gate that drops gap-fill zeros, flat runs, clipping and glitch spikes; three years of XCache logs explored before modelling.
- 05
Run experiments & ablations
Change one thing at a time and see what actually carries the result.
- 5-seed ensembles for stability; a nearest-single-station ablation (R² 0.84 → 0.42) that proves the graph fusion matters.
- A lookback sweep across forecast horizons, and an ablation on only the cleanest inputs for the inverse-FEA model.
- 06
Analyze honestly
Use metrics that can't be gamed, and write down the limits.
- ROC-AUC and MCC on imbalanced data instead of accuracy; MCC instead of recall for alerts.
- Known limits stated plainly, e.g. the magnitude model under-predicts the very largest events.
- 07
Conclude & communicate
Say exactly what the evidence supports, to the people who need it.
- The inverse-FEA ceiling is a data limit: written up for the team's IEEE paper, and it redirected their effort.
- SeismicSoCal's results published next to their baselines on the live site.
- 08
Iterate
Every answer sets up the next question.
- Test a ~120-day lookback at the 30-day horizon.
- Wire the early-warning model into live alerts and extend coverage statewide.
Step 8 feeds straight back into step 1.