In my previous post I worked out how a negative carrier test for an autosomal recessive disorder updates a sibling's risk. Not much, as it turns out, unless the background rate is already quite high. ApoE is more interesting because it involves three common alleles.
ApoE in a nutshell
The APOE gene codes for a protein involved in lipid transport and cholesterol metabolism. It's the principal cholesterol carrier in the brain and interacts with the low-density lipoprotein receptor to clear lipoprotein particles from circulation. It comes in three common variants, called ε2, ε3, and ε4. Since you inherit one copy from each parent, there are six possible genotypes: ε2/ε2, ε2/ε3, ε2/ε4, ε3/ε3, ε3/ε4, and ε4/ε4.
There are a few reasons to care about this genotype: Alzheimer's risk, cardiovascular disease, and immunity.
Alzheimer's risk
ε4 is the strongest common genetic risk factor for late-onset Alzheimer's disease. The table below gives odds ratios for AD relative to ε3/ε3 in European-ancestry populations drawing data from a modern 337,484-person UK Biobank analysis (Lumsden et al. 2020, white British participants; per-genotype values from appendix table S4)1 alongside the classic Farrer et al. (1997) meta-analysis2 (Caucasian; 5,930 AD patients, 8,607 controls across 40 research groups).
| Genotype | UK Biobank, Lumsden 20201 (95% CI) | Farrer 19972 (95% CI) |
|---|---|---|
| ε2/ε2 | 0.46 (0.06–3.27) | 0.6 (0.2–2.0) |
| ε2/ε3 | 0.77 (0.53–1.11) | 0.6 (0.5–0.8) |
| ε2/ε4 | 1.83 (1.10–3.05) | 2.6 (1.6–4.0) |
| ε3/ε3 | 1.00 (reference) | 1.0 (reference) |
| ε3/ε4 | 3.69 (3.08–4.40) | 3.2 (2.8–3.8) |
| ε4/ε4 | 13.5 (10.6–17.2) | 14.9 (10.8–20.6) |
The two sources concord closely. Only the ε3/ε4 and ε4/ε4 associations cleared the phenome-wide–corrected significance threshold of this UK Biobank study; the protective ε2 genotypes are too sparse in a healthy middle-aged cohort, leaving the wide confidence intervals seen above.
AD effect sizes vary substantially by ancestry: Farrer reports the ε4/ε4 odds ratio ranging from 33.1 in Japanese and 14.9 in Caucasian cohorts down to 5.7 in African Americans and just 2.2 in Hispanics2.
ε4/ε4 individuals are 88% AD-biomarker-positive by age 803. Taking ε2/ε2 as the reference genotype, Williams et al. (2026) estimate that ε3 and ε4 together account for the vast majority of AD with attributable fractions ranging from ~72% (FinnGen) to ~93% (ADGC) across cohorts4. That is, the ε3/ε3 "baseline" is neutral only by convention, not in absolute terms.
Belloy et al. (2019)5 has another synthesis.
Cardiovascular risk
ApoE also has an impact on cardiovascular health via LDL cholesterol. ε2 is also favorable here: the ε2 isoform lowers circulating LDL while ε4 raises it. LDL-C, and coronary risk it associates with, climbs approximately linearly across the genotypes from ε2/ε2 to ε4/ε4. In the Bennet et al. (2007) meta-analysis (86,067 people for lipids; 37,850 coronary cases), the coronary heart disease (CHD) odds ratios relative to ε3/ε3 are6:
| Genotype | CHD OR (95% CI) |
|---|---|
| ε2/ε2 | 0.83 (0.55–1.25) |
| ε2/ε3 | 0.82 (0.72–0.92) |
| ε2/ε4 | 0.93 (0.81–1.07) |
| ε3/ε3 | 1.00 (reference) |
| ε3/ε4 | 1.05 (0.99–1.12) |
| ε4/ε4 | 1.22 (1.08–1.38) |
Pooled by carrier status, ε2 carriers had roughly 20% lower coronary risk than ε3/ε3 (OR 0.80, 95% CI 0.70–0.90), while ε4 carriers were not significantly different (OR 1.06, 0.98–1.15)6. Coronary disease is not the whole cardiovascular story, though: independent of its modest coronary effect, the ε4 allele raises susceptibility to stroke and cerebrovascular disease (an allele-level effect, so it applies to both ε3/ε4 and ε4/ε4)5.
Modern direct-imaging data reinforce ε2's benefit: in the PESA cohort of 3,887 asymptomatic midlife adults, ε2 carriers had markedly less subclinical atherosclerosis across the carotid, femoral, and coronary beds (coronary odds ratio 0.53), largely independent of their lower LDL7. In Haffner et al.'s biethnic population study, LDL peak particle size decreased stepwise from ε2/ε3 → ε3/ε3 → ε3/ε4, leaving ε2 carriers prone to pattern-A LDL (large and buoyant), while ε4 carriers had smaller, denser pattern-B particles8. This is a well-established feature of the APOE–LDL-subclass relationship9.
The above looks good for ε2, but there's a problem. The same mechanism that causes it to lower LDL also makes it clear chylomicron and VLDL remnants poorly, so ε2 tends to raise triglycerides. In particular, ~5% of ε2/ε2 individuals develop type III hyperlipoproteinemia10, and even then usually only with a second insult like insulin resistance or hypothyroidism11. Elevated triglycerides drive pattern-B LDL, which is atherogenic and bad.
How can we reconcile ε2 being prone to both pattern A and pattern B? Preexisting factors. In an ε2 carrier who is already hypertriglyceridemic or has established coronary disease, the balance flips toward pattern B12. In the PESA cohort, ε2's atheroprotection vanished in carriers with triglycerides ≥150 mg/dL7. So ε2's cardiovascular benefit is real but triglyceride-contingent: it holds as long as triglycerides stay in range.
ApoE and immunity
ApoE also has immune effects. ApoE-null mice are markedly more susceptible to Listeria and Klebsiella infection10. In contrast, apoE facilitates herpes simplex (HSV), HIV, and dengue entering cells, and ε4 facilitates this the most.13. ApoE binds cell-surface glycosaminoglycans in the rank order E4 > E3 > E2, and that ordering tracks how readily each isoform helps HSV-1 attach to and enter cells13. Clinically, ε4 carriers get more cold sores: one study linked the ε4 allele to oral herpetic lesions (relative risk ≈4.6) while finding no effect on how often the virus reactivated or shed14. The same pattern holds for HIV: apoE4 is a weaker inhibitor of viral membrane fusion than apoE3, so it enhances HIV cell entry. In a 1,267-patient cohort, ε4/ε4 homozygotes progressed faster and carried higher steady-state viral loads in a dose-dependent way (ε4/ε4 > ε4/non-ε4 > non-ε4), and a separate cohort found ε4 carriers reaching HIV disease in 8.3 versus 10.5 years even though APOE genotype did not change the risk of acquiring HIV in the first place (ε2 = ε3 = ε4)15.
This ties back to Alzheimer's. ε4 is over-represented among AD patients who carry HSV-1 but not among those who don't. This suggests some of ε4's Alzheimer's risk may run through viral susceptibility10. For an ε2/ε3 result like mine, the E4 > E3 > E2 gradient means my genotype is among the least facilitating on this infectious-disease axis15.
Miscellaneous facts
- Nearly all other mammals, and all the great apes, carry an ε4-like apoE. ε3 and ε2 arose more recently in human evolution, each by a single-nucleotide change10.
- In a study of 40-year-old Danish men, ε3/ε3 men averaged 1.93 children with only 6% childless, versus 1.50 children (26% childless) for ε4 carriers and 1.66 (19% childless) for ε2 carriers. This is a statistically significant fertility edge for ε3/ε3 (P = 0.003)16.
- The frequency of ε4 is lowest around the historical cradles of agriculture. These are thought to be among the first places dense enough to sustain epidemic viral disease and reinforce ApoE's role in immunity10.
An answer to this question on Stack Overflow.
Question
PyTorch contains these macros which use __builtin_expect:
#define C10_LIKELY(expr) (__builtin_expect(static_cast<bool>(expr), 1))
#define C10_UNLIKELY(expr) (__builtin_expect(static_cast<bool>(expr), 0))
Apart from the fact that __builtin_expect is compiler specific,
what is the difference between these and C++20's [[likely]] and [[unlikely]] attributes?
Answer
Quick Answer: Expressivity
The difference is in the degree of expressivity. With __builtin_expect we can write:
if (LIKELY(x()) && UNLIKELY(y())) {
while with [[likely]] we can only write
if (x && y) [[likely]] {
Does it matter?
Yes.
Let's construct two programs to explore the difference.
Using builtin
// builtin.cpp
// Method A: per-operand hints via __builtin_expect.
// Opaque predicates + sinks so nothing is folded away and both
// short-circuit branches must be emitted as real branches.
extern bool x();
extern bool y();
extern void body();
extern void keep_going();
#define LIKELY(expr) (__builtin_expect(static_cast<bool>(expr), 1))
#define UNLIKELY(expr) (__builtin_expect(static_cast<bool>(expr), 0))
// "x almost always true, y almost always false"
void hint() {
#ifdef LIKELY_FIRST // For exploring orderings
if (LIKELY(x()) && UNLIKELY(y())) {
#else
if (UNLIKELY(x()) && LIKELY(y())) {
#endif
body();
}
keep_going();
}
// Control: single condition.
void single() {
if (LIKELY(x())) { body(); }
keep_going();
}
Using [[likely]]
// likely.cpp
// Method B: statement-level hint via the standard [[likely]]/[[unlikely]] attribute.
// Opaque predicates + sinks so nothing is folded away and both
// short-circuit branches must be emitted as real branches.
extern bool x();
extern bool y();
extern void body();
extern void keep_going();
// only says "the taken branch (body) is unlikely"
void hint() {
if (x() && y()) [[unlikely]] {
body();
}
keep_going();
}
// Control: single condition.
void single() {
if (x()) [[likely]] { body(); }
keep_going();
}
Compiling and comparing
We can compile to assembly
clang++ -O2 -std=c++20 -S -o builtin_likely_first.s builtin.cpp -DLIKELY_FI
RST
clang++ -O2 -std=c++20 -S -o builtin_unlikely_first.s builtin.cpp
clang++ -O2 -std=c++20 -S -o likely.s likely.cpp
when we compare
diff builtin_likely_first.s likely.s
There's not a meaningful difference.
However, when we compare builtin_unlikely_first, we find the following:
# diff -y --width=80 --expand-tabs builtin_unlikely_first.s likely.s
# LEFT = UNLIKELY(x()) && LIKELY(y()) RIGHT = x() && y() [[unlikely]]
# gutter: | changed < left only > right only (blank = same)
.file "builtin.cpp" | .file "likely.cpp"
.text .text
.globl _Z4hintv .globl _Z4hintv
.p2align 4 .p2align 4
.type _Z4hintv,@function .type _Z4hintv,@function
_Z4hintv: _Z4hintv:
.cfi_startproc .cfi_startproc
# %bb.0: # %bb.0:
pushq %rax pushq %rax
.cfi_def_cfa_offset 16 .cfi_def_cfa_offset 16
callq _Z1xv callq _Z1xv
testb %al, %al testb %al, %al
jne .LBB0_1 | je .LBB0_3
> # %bb.1:
> callq _Z1yv
> testb %al, %al
> jne .LBB0_2
.LBB0_3: .LBB0_3:
popq %rax popq %rax
.cfi_def_cfa_offset 8 .cfi_def_cfa_offset 8
jmp _Z10keep_goingv jmp _Z10keep_goingv
.LBB0_1: | .LBB0_2:
.cfi_def_cfa_offset 16 .cfi_def_cfa_offset 16
callq _Z1yv <
testb %al, %al <
je .LBB0_3 <
# %bb.2: <
callq _Z4bodyv callq _Z4bodyv
popq %rax popq %rax
.cfi_def_cfa_offset 8 .cfi_def_cfa_offset 8
jmp _Z10keep_goingv jmp _Z10keep_goingv
.Lfunc_end0: .Lfunc_end0:
.size _Z4hintv, .Lfunc_end0- .size _Z4hintv, .Lfunc_end0-
.cfi_endproc .cfi_endproc
.globl _Z6singlev .globl _Z6singlev
.p2align 4 .p2align 4
.type _Z6singlev,@function .type _Z6singlev,@function
_Z6singlev: _Z6singlev:
.cfi_startproc .cfi_startproc
# %bb.0: # %bb.0:
pushq %rax pushq %rax
.cfi_def_cfa_offset 16 .cfi_def_cfa_offset 16
callq _Z1xv callq _Z1xv
testb %al, %al testb %al, %al
je .LBB1_2 je .LBB1_2
# %bb.1: # %bb.1:
callq _Z4bodyv callq _Z4bodyv
.LBB1_2: .LBB1_2:
popq %rax popq %rax
.cfi_def_cfa_offset 8 .cfi_def_cfa_offset 8
jmp _Z10keep_goingv jmp _Z10keep_goingv
.Lfunc_end1: .Lfunc_end1:
.size _Z6singlev, .Lfunc_end .size _Z6singlev, .Lfunc_end
.cfi_endproc .cfi_endproc
.ident "clang version 21.1.8 .ident "clang version 21.1.8
.section ".note.GNU-sta .section ".note.GNU-sta
.addrsig .addrsig
Let's focus on the key changes. This section
testb %al, %al testb %al, %al
jne .LBB0_1 | je .LBB0_3
> # %bb.1:
> callq _Z1yv
> testb %al, %al
> jne .LBB0_2
shows that after testing x(), the two versions part ways.
On the left (UNLIKELY(x()) first), the whole fast path is the single instruction jne .LBB0_1. Since x is expected to be false, "x is false" just falls straight through
toward keep_going(), and the code only jumps away, to a separate block, in the rare case that x is true. y() isn't here at all (it's been moved farther down).
On the right ([[unlikely]]), after testing x, y() is also called and tested (callq _Z1yv -> testb), so y() stays on the path.
The next section is
.LBB0_1: | .LBB0_2:
.cfi_def_cfa_offset 16 .cfi_def_cfa_offset 16
callq _Z1yv <
testb %al, %al <
je .LBB0_3 <
# %bb.2: <
On the left we have the section the previous left-side section jumped to. y() is compared here in the rare case that's necessary (callq _Z1yv -> testb.
On the right the matching block has none of those lines because y() was already evaluated further up.
What difference does it make?
Modern CPUs don't just execute instructions one at a time. They have many layers of complexity to improve code performance and these layers interact with how we write code.
The instruction cache and the front-end. The CPU fetches instructions in contiguous chunks (cache lines, ~64 bytes). Instructions that sit next to each other in memory get fetched together "for free." Placing the unlikely y() and body() in a separate block way at the end of the function means that the most likely path of execution is pulled into cache and then used. In contrast, if unlikely y() follows likely x() directly then cold code evicts the useful hot code. Rearranging code so this doesn't happen is called hot/cold splitting and means fewer cache lines are touched on the hot path and better use is made of the instruction cache.
Branch prediction and fall-through. The CPU attempts to predict whether a branch will be taken or not (so it can prefetch and do speculative execution). If the CPU predicts wrong this work is wasted and a full pipeline flush may be necessary.
Compilers arrange code so the expected path is the fall-through (no jump) and the unexpected path is the one that jumps away. Note that putting UNLIKELY above meant that the hot path has one fewer branch prediction then the when we used [[unlikely]]. That gives the CPU a better shot at getting things right, in addition to the instruction cache benefits.
As a caveat, a person can really overthink this. You'd only notice the difference between these in a really hot path. If this is something you actually worry about it would be better to use profile-guided optimization (PGO): the compiler can measure actual branch frequencies and reorganize code accordingly, ignoring hints. This is much better than using hints. So hints matter most in situations where we can't profile and suspect a hot path.
Branch-hint layout matrix (clang 21, -O2)
We can explore this for all orderings and logical operators by evaluating
void f() {
if ( <cond over x() and y()> ) { body(); }
keep_going();
}
And we find that && and || will short-circuit. In contrast, ^ never short-ciruits (both x() and y()) always run, so per-operand hints can't steer it.
Legend: "body usually runs taken?" = is the body the common case? "cold tail" = pushed to an
out-of-line block. "inline" = on the straight-line fast path. "always" = the
body() call is emitted unconditionally (no short-circuit).
| Expression | body usually runs? | what happens at the x test |
body() |
y() |
|---|---|---|---|---|
LIKELY(x) && UNLIKELY(y) |
no | je away if x false; fall through to eval y(); jne to body |
cold tail | inline |
UNLIKELY(x) && LIKELY(y) |
no | jne away to eval y()+body(); fall through to keep_going() |
cold tail | cold tail |
x && y [[unlikely]] |
no | je away if x false; fall through to eval y(); jne to body |
cold tail | inline |
x && y [[likely]] |
yes | je away if x false; fall through to eval y(); body() inline |
inline | inline |
LIKELY(x) \|\| UNLIKELY(y) |
yes | je away to eval y(); fall through to body() inline |
inline | cold tail |
UNLIKELY(x) \|\| LIKELY(y) |
yes | jne away to body(); fall through to eval y() |
inline | inline |
x \|\| y [[likely]] |
yes | jne away to body(); fall through to eval y() |
inline | inline |
x \|\| y [[unlikely]] |
no | jne away to body() (cold); fall through to eval y(); fall to keep_going() |
cold tail | inline |
LIKELY(x) ^ UNLIKELY(y) |
hints ignored* | eval x(), eval y(), cmpb; je skip; body() inline |
inline | always |
UNLIKELY(x) ^ LIKELY(y) |
hints ignored* | eval x(), eval y(), cmpb; je skip; body() inline |
inline | always |
x ^ y [[likely]] |
yes | eval x(), eval y(), cmpb; je skip; body() inline |
inline | always |
x ^ y [[unlikely]] |
no | eval x(), eval y(), cmpb; jne to body() (cold); fall to keep_going() |
cold tail | always |
An answer to this question on Stack Overflow.
Question
I am not sure Stack Overflow is the right place for this question but I'm shooting my shot.
I have a 2D field sampled on a global longitude-latitude structured curvilinear grid and I would like to compute how much of the field's variance/energy is associated with different spatial scales.
The data is stored as a NumPy array of shape (nt, nx, ny). For each grid point I also have its longitude and latitude coordinates as 2D arrays.
The grid is regular in longitude and latitude, for example a constant angular spacing of 0.25°, but it is not regular in physical distance. Due to meridian convergence, two neighboring points separated by 0.25° longitude are much closer together near the poles than near the equator. I also have the physical area of each grid cell and local scale factors relating angular distances to physical distances.
My initial approach was to compute a 2D FFT and derive a power spectrum. However, a standard FFT assumes uniform sampling, which is not true in physical space for a longitude-latitude grid. I then tried interpolating the data onto a more uniform grid before applying the FFT, but the interpolation noticeably smooths small-scale features and changes the resulting spectrum.
So my question is, what is the standard way to perform a scale decomposition in this situation?
Can I apply a standard FFT directly in index space if the goal is to compare multiple fields defined on the same grid? Should I look into non-uniform FFTs applicable when the sampling locations are known through longitude and latitude coordinates? Should I look at an alternative approach altogether?
Answer
Perhaps convert the points to 3D (x,y,z) or use a spatial binning (eg, https://github.com/r-barnes/dggridr/).
tl;dr There are a bunch of planning tools near the bottom of the post.
Tomales Bay is one of the best places in California to kayak through bioluminescent plankton, but you can only see it if three things align: it has to be dark enough (moon below the horizon), late enough (after full darkness), and the right season (late spring through fall when dinoflagellate populations peak). What follows is a trip report followed by some planning tools for your own trip.

The First Trip — September 23–24, 2022
We reserved Boat Site B at Point Reyes National Seashore Campground, a group site on the west shore of Tomales Bay at Marshall Beach, accessible only by water.

The original plan was to park overnight at, and launch from, Miller Boat Launch on the east shore and paddle across to camp.

Straight-line, the route is about 1.78 miles.

By timing the crossing for slack tide, though, we could have cut straight across the Bay — only about 4,500 feet — and then head southeast along the shore to the camp site. This would have minimized exposure to open water risks.

But, in planning this, I chickened out a bit. The group included people with variable kayaking experience and high afternoon winds can lead to large waves on the Bay, so we instead launched from Chicken Ranch Beach to the south, and paddled 5.29 miles along the shore so that there'd be an easy way to bail if things went wrong. This choice brought us into Type II Fun territory.




I chose the date to maximize darkness. Moonset on September 23 was at 6:19 PM — well before sunset — leaving the whole evening dark. (The planner below includes all this information in an easy-to-use form.)

In addition, civil twilight ended at 7:39 PM, nautical twilight at 8:06 PM, and astronomical twilight at 8:33 PM.

Winds pick up in the afternoons so launching before noon is recommended and, in fact, local outfitters won't rent kayaks after noon even though winds tend to drop off again later in the day. Given the tide and light, starting around 4:00 PM during slack tide should have given enough time to paddle and set up camp before dark.
However, due to folks' work schedules, we weren't able to get onto the water until 5:12 PM (a vanguard from our group acquired the kayaks). This meant meant fighting the incoming tide that started at 5:53 PM on the 23rd. But not just the tide! Also that wind. The result was a brutal, soaking paddle. We got to camp at 7:17 PM, so the trip took about two hours.
On the way out we didn't time it much better and got on the water at 11:25 AM to fight the outgoing tide on our way back. The winds were calm and the paddling was much easier than the day before. We took it slow and arrived at 1:18 PM, so the trip took about two hours.

Most beaches on the west shore are tidally dependent and will disappear at tides above 5 ft. Since September 23 had a high tide of 5.28 ft at 11:11 PM, we needed to set up camp well above the waterline.

Cold and wet, some members of the trip retreated to tents immediately to get warm and no one was up for venturing back out again into the dark. The fabled bioluminescence didn't show itself in the bay near our tents.
But the weekend wasn't a wash. Our visit corresponded with the annual campout of the Traditional Small Craft Association. They played folk music by their fire well into the night.



The next morning we went hiking up and around our mini-bay.


It took a while, so we investigated climbing across the cliffs to get back to camp.

But were spared that adventure by the TCSA coming over in one of their boats!




The Second Trip — October 6–7, 2023
I returned a year later with a friend to try again. Blue Waters normally rents sit-on-top kayaks, but as a former kayak guide I was able to get a lighter enclosed double. This time we used the Miller Boat Launch. The winds were, by luck, exceptionally calm when we got on the water at 3:30 PM. By 4:39 PM we'd successfully crossed to the opposite shore and were making our way along it and by 5:47 PM we had our tent set up at Tomales Beach (which is closer to Miller than Marshall Beach).




After dark, we got back on the water and kayaked from Tomales Beach up to White Gulch Beach - which is in a deep bay directly west of Hog Island. We found bioluminescence the whole way, and it was incredible. Every paddle stroke lit up the waters and disturbed kelp shot lightning bolts away from us.

An answer to this question on Stack Overflow.
Question
I'm not a specialist, but as far as I know, a bit of information in a QR-code is coded more than once, and it is defined as the redundancy level
How can I estimate a QR-code redundancy level ? Is where an mobile app or a website where I can test my QR-code redundancy level easily ? If not, is it an easy algorithm that I can implement ?
Redundancy is sorted in different categories according to this website, but I'd like to have the direct percentage value if possible
Answer
QR codes contain a couple of bits which indicate the error correction level, as depicted below (source):
