Skip to content

Source

engines/quan_loncar_meep_lab/notes/device23_seed_lineage.md · assembled 2026-09-05 15:22 UTC.


Device23 launch seed: preserve the device before optimizing it

The inherited design23_20260801/undercut_mirror.json is useful as a 1550 nm layer-stack control, but it is not Device23. It is a 900 x 220 nm undercut beam with eleven identical holes, including one at x = 0; it has no defect and its measured localization was 0.21.

The supplied Device23 fabrication notebook instead defines an oxide-backed, 795.692787 nm cavity in a 556.263652 x 300 nm SiN beam. It has 42 positive holes reflected across x = 0 (84 total), no central hole, a 234.625430 nm pitch, 22 tapered pairs, and 20 mirror pairs. The recovered FDTD result is Q = 11,812.391. That exact lineage control is exported as benchmarks/design23_20260801/device23_canonical_84hole.json.

The full article is not the practical local launch seed. A later Device23 campaign retained the inner 20 pairs and physically evaluated its accepted step 6 in Tidy3D. The checked result is:

observable qualified value
wavelength 770.029669 nm
Q 1,113.003
anthracene-normalized V 6.180184 (lambda/n)^3
Q/V 180.092
two-port TE0 beta, power 0.878603
two-port TE0 beta, Q cross-check 0.886978
TE0 fraction of x-directed flux 0.962033
requested/reported lead modes 4 (not exhaustive)

All twelve coherent-observable checks pass, including incoming-wave rejection, power closure, on-resonance alignment, TE polarization, and the independent power/Q beta cross-check. The coherent Tidy3D simulation requested four modes, so its four reported effective indices are provenance, not proof that the lead has exactly four guided modes; local qualification must search for additional weak branches. Its exact 62-coordinate geometry is exported as benchmarks/design23_20260801/device23_40hole_fdtd_seed.json; this is the free-form campaign seed. The complete source is engines/design23_v1/runs/20260728T013605Z_design23_40hole_composite_bfgs/ fdtd_step_0006_diagnostic/diagnostic_result.json.

Stack translation, not a stack change

The legacy coordinates are translated upward by 150 nm for the new LayerStack convention:

material legacy z (um) layered-engine z (um)
SiN [-0.15, 0.15] [0.0, 0.3]
anthracene [0.15, 0.35] [0.3, 0.5]
PVA [0.35, 0.55] [0.5, 0.7]
anthracene emitter 0.25 0.4

The oxide substrate is retained. It is physically semi-infinite; the JSON bottom_um = -0.6 is only a numerical truncation, and the stack builder extends the bottommost unetched substrate to the lower physical/PML boundary. The launch seed is deliberately not undercut because no undercut version of this geometry has equivalent pole and modal evidence.

Exact 770 nm native lead certificate

The substrate-backed step-6 lead was qualified locally at its exact 770.029669 nm target, independently of the unrelated 1550 nm undercut control. The mode solver uses a full-y/full-z vector Yee Bloch cross-section and a one-plane Lorentz-reciprocity E/H projection. A 25/um solve with a 20/um coarse control, 24 beta samples, 30 eigenpairs, and deep port padding of 0.6 um in y and 1.0 um in z produced the following evidence:

control qualified value
stable localized branch count 6
selected TE0 effective index 1.778786891
selected TE0 Ey fraction 0.909672346
target beta grid difference 0.019189852
target boundary energy 0.000051042
target reciprocity self error 4.44e-16
opposite-direction leakage 4.30e-17
maximum other-branch overlap 2.57e-10

The selected TE0 is branch 1 rather than the largest-beta branch because branch 0 is TM-like. The historical Tidy3D monitor requested exactly four modes, so the six local branches do not contradict that capped report.

The deeper two-dimensional mode box is deliberately larger than the pilot's three-dimensional physical cross-section. Before sampling the 3-D field, the TE0 functional is cropped to y = +/-0.778131826 um and z = [-1.05, 1.15] um. Its retained self-overlap is 0.999913692; the omitted 8.63078e-5 is below the pinned 1e-3 gate, after which the retained functional is renormalized to unit target overlap. Sampling therefore never enters the transverse PML.

Weak unwanted branches close to the substrate light line have 7--9% boundary energy even in the deep box. This is a recorded nonfatal warning, not a numerator qualification failure: the hard numerator projects only the well-converged TE0, and the physical full-y dipole-work denominator counts all other guided and radiative power as unwanted. Fine/coarse branch-count stability remains mandatory so a missing branch cannot be hidden by that policy. The complete 20/16 failed control and passing 25/20 follow-up are in benchmarks/design23_20260801/device23_770_layered_port_control.json.

Geometry and launch gates

On the 4 nm level-set grid the raw seed has 40 voids and a 77.695 nm minimum feature. The production 12 nm wall filter, one-cell step filter, reinitialization, and 80 nm island-area pruning retain all 40 voids; the settled field has a 76.878 nm minimum feature, 109.810 nm longitudinal gap, and 146.364 nm sidewall web. Therefore the free-form feature floor is 70 nm. The 80 nm island setting is safe because it is an area-equivalent threshold, not a minor-diameter threshold.

The default adaptive 96-plus-node spline is too expensive on forty level-set voids. A fixed-node audit against 2,048 samples per void is recorded in benchmarks/design23_20260801/device23_aperture_resampling.json. At 32 nodes, the conservative pre-spline polyline has a 1.790 nm maximum Hausdorff error and 0.867% maximum area error. At 48 nodes those fall to 0.987 nm and 0.412%, while the 76.878 nm minimum feature and 109.810 nm gap are unchanged. The named Device23 pilot fidelity therefore uses 48 fixed nodes per void and disables only the adaptive node multiplier; the global 96-node default is unchanged. This aperture approximation still needs a 32-versus-48 pole comparison before it can become a certification fidelity.

The resulting fixed-48, all-order-one mesh has 619,474 degrees of freedom and builds its mesh/space in 143.56 seconds under concurrent eigensolver load; the machine-readable record is device23_pilot_mesh.json. In contrast, the named medium mixed-order mesh reaches 2,072,871 unknowns even with exact OCC ellipses and is refused by the 500k production safety gate before eigenanalysis. This is why the campaign labels the tractable setting device23_pilot_p1 and does not present it as an order-converged Q/V result.

The corresponding physical full-y point-source domain was also built without assembly or factorization. At the same fixed-48, all-order-one fidelity it has 1,252,734 degrees of freedom, builds in 70.27 seconds, and uses 722,956 kB peak RSS without swap. This expected near-doubling is recorded separately in device23_pilot_full_y_mesh.json; it is a resource measurement, not an LDOS, volume, or coupling result.

That pilot then failed the pole smoke. With the exact OCC ellipses (so aperture resampling was not involved), the closest candidate was 1.2938 - 0.0085i /um, residual 4.3e-6 and legacy central-patterned-energy selection value 0.176. None of eight candidates met the 1e-8 residual gate; the largest legacy selection value was 0.252. Those values are not comparable to the later Ey-monitor localization. The result is recorded in device23_40hole_stack_pole_smoke.json. Consequently device23_pilot_p1 is a mesh-scaling diagnostic only and is not safe to launch as a Q/V or coupling-restoration optimizer fidelity.

Increasing the Arnoldi request from 8 to 16 candidates repaired convergence but not the fidelity. On the settled fixed-48 field it returned residual 1.14e-13 at 765.439 nm and Q = 120.64. Its reported 0.2635 selection value is the legacy central-patterned-energy fraction because the process imported the solver before the Ey-monitor selector update; it must not be compared to the current Ey localization metric. The run used 619,474 DOF, 342 seconds, and 17,562,636 kB peak RSS without swap. Since the coherent FDTD seed has Q = 1,113, this converged Q = 120 branch still fails the fidelity gate and cannot launch the optimizer. The full record is device23_40hole_stack_pole_p1_c16.json.

A selective-order mesh ladder on the same fixed-48 settled field gives:

half-y fidelity DOF
base/PML order 1, cavity order 2 for abs(x) < 2.8 um 1,447,138
base/PML/cavity order 1, inner order 2 for abs(x) < 1.4 um 972,976
base order 1, PML order 2, cavity order 2 for abs(x) < 1.4 um 1,004,130

The last case is the smallest ladder point that raises both the central cavity and the radiation-sensitive PML while staying near one million unknowns. It is the next scientifically defensible pole fidelity if a larger all-order-one Arnoldi space still cannot recover the branch. The exact settings and mesh times are stored in device23_order_ladder.json.

The seed begins below the requested 95% two-port TE0 constraint. It must enter the optimizer through the measured restoration phase, not be mislabeled as feasible. Launch remains fail-closed until the local layered pole and E/H feedthrough projections reproduce a localized branch and a physical modal budget. Raw x-PML absorption cannot satisfy that gate.

Channel resource envelope, 2026-08-02

The two-port feedthrough is measured on a full-y domain because a point current on a PEC mirror has no stable H(curl) normal trace. That domain is 1,252,734 unknowns against the pole's 619,474, and it is driven rather than an eigenproblem, so it factorizes where the pole iterates.

Measured, not estimated:

solver limit outcome
UMFPACK 38 GiB virtual numeric factorization out of memory
MKL Pardiso 36 GiB resident OOM killed at 36 G peak, 15m54s wall, 40m CPU

Pardiso is the better tool -- multithreaded, far less fill -- and it still does not fit in 36 GiB on a 45 GiB machine at the pilot channel fidelity.

Two practical notes for anyone repeating this. ulimit -v is the wrong guard: it caps virtual address space, and MKL with its OpenMP arenas reserves generously enough that Pardiso reports out-of-memory while resident use is far below the cap. A cgroup limit through systemd-run --user --scope -p MemoryMax=... -p MemorySwapMax=0 caps resident memory instead, which is the quantity that actually threatens the machine, and it kills only its own scope. And installing MKL reaches further than the call sites that name a solver: ArnoldiSolver selects the default sparse solver itself, so the pole solve moved to Pardiso without being asked.

Python buffers stdout to a pipe, so a run killed mid-solve reports nothing at all unless invoked with -u. The first attempt lost its completed pole result that way.