2026-06-23
We reason in an informal metatheory (ZFC). To avoid proper-class pathologies we fix once and for all a set $G$, the global domain of all things, rather than a proper class (Remark 7). All cardinal arithmetic is metatheoretic. “$\aleph_0$” denotes countable infinity; “at most countable” means finite or countably infinite.
Definition 1.1 [Formal language]. A formal language is a pair $(\Sigma, L)$ where the alphabet $\Sigma$ is an at most countable set of symbols and $L \subseteq \Sigma^{*}$ is a set of finite strings over $\Sigma$ (the well-formed expressions). Since $\Sigma$ is at most countable, so is $\Sigma^{*}$, hence $|L| \le \aleph_0$.
Definition 1.2 [Representation system]. A representation system over $G$ is a triple $\mathcal{S} = (\Sigma, L, \rho)$ where $(\Sigma, L)$ is a formal language and the denotation $\rho \colon L \rightharpoonup G$ is a partial function. read $\rho(w) = x$ as “the expression $w$ represents the thing $x$.” Functionality encodes that a fixed expression, in a fixed system, picks out at most one thing; ambiguity is handled in Remark 1.
Definition 1.3 (The representable and the irrepresentable). Fix an at most countable family of representation systems $\mathcal{R} = \{\mathcal{S}_i\}_{i \in I}$, $|I| \le \aleph_0$, $\mathcal{S}_i = (\Sigma_i, L_i, \rho_i)$. Define
\[ O_f \;:=\; \bigcup_{i \in I} \operatorname{ran}(\rho_i) \;=\; \{\, x \in G : \exists\, i\in I,\ \exists\, w \in L_i,\ \rho_i(w)=x \,\}, \qquad O_n \;:=\; G \setminus O_f . \]$O_f$ is the class of formally representable things, $O_n$ its complement in $G$, and $G = O_f \sqcup O_n$ is a partition.
Assumption A (Verbalizable $\iff$ Formally Representable). A thing is verbalizable iff it lies in $O_f$ for the family $\mathcal{R}$ comprising the formalizable fragments of natural language. Assumption A is what licenses transporting any conclusion about $O_f / O_n$ into a conclusion about what can and cannot be said. I guess its an axiom, not a theorem. see §6 for plausibility and the issue raised in Q2. I guess in short, this concerns finished utterances, not the open ended act of meaning-making.
Terminology note (information $:=$ $O_f$). We deliberately speak of “things,” not “information.” In Shannon and Kolmogorov's theories “information” is formal by construction (bits, codewords, programs), so naming the elements of $G$ “information” would would predicate $O_n = \emptyset$. So: \[ \text{information} \;:=\; O_f \quad(\text{the finitely codeable part of }G), \]
In general, things $\supsetneq$ information, implying that $O_n$ is whatever is "non-informational" (See Q7.)
Lemma 2.1 (Representability is countable). $|O_f| \le \aleph_0$.
Proof. For each $i$, $L_i$ is at most countable (Def. 1.1) and $\rho_i$ is a partial function, so $|\operatorname{ran}(\rho_i)| \le |L_i| \le \aleph_0$. Then $O_f$ is a union indexed by the at most countable $I$ of at most countable sets; by the countable union theorem (countable choice suffices), $|O_f| \le \aleph_0 \cdot \aleph_0 = \aleph_0$. ∎
Remark 1 (ambiguity sucks). If denotation is relational, one expression $w$ admitting a set $\rho_i[w] \subseteq G$ of referents, the lemma survives provided each $\rho_i[w]$ is at most countable, since $O_f$ is then still a countable union of countable sets. It fails only if a single finite expression may coorespond to uncountably many things at once; but any system with a word or phrase like that can contain no information, and trivializes Assumption A.
Remark 2 (description-length / compression form). Define the description length $\ell(x) := \min\{\,|w| : \rho_i(w)=x \text{ for some } i\,\}$, with $\ell(x)=\infty$ when $x$ has no representative; equivalently $\ell$ is a Kolmogorov complexity relative to $\mathcal{R}$. Then $O_f = \{x : \ell(x) < \infty\}$. For each $n$ there are at most $\sum_{k \le n}\sum_i |\Sigma_i|^{k}$ expressions of length $\le n$, a finite number per fixed finite alphabet (and countable overall), so $\{x : \ell(x) \le n\}$ is at most countable and $O_f = \bigcup_{n} \{x : \ell(x) \le n\}$ is countable. This re-proves Lemma 2.1 using compression: to represent is to compress to a finite codeword, and only countably many finite codewords exist.
Proposition 3.1. If $|G| > \aleph_0$, then $O_n \neq \emptyset$; more precisely $|O_n| = |G|$ while $|O_f| \le \aleph_0$. Both $O_f$ and $O_n$ are proper and nonempty as soon as $G$ contains at least one representable thing.
Proof. By Lemma 2.1, $|O_f| \le \aleph_0 < |G|$. If $O_n$ were at most countable then $|G| = |O_f \cup O_n| \le \aleph_0 + \aleph_0 = \aleph_0$, a contradiction; so $|O_n| > \aleph_0$, hence $O_n \ne \emptyset$. Moreover $|O_n| \le |G| = |O_f \cup O_n| \le |O_f| + |O_n| = |O_n|$ (the last step since $|O_n|$ is infinite and absorbs $|O_f| \le \aleph_0$), so $|O_n| = |G|$ by Cantor–Schröder–Bernstein. $O_f \neq \emptyset$ exactly when some $\rho_i$ is nonempty. ∎
Proposition 3.2 (Stochastism changes nothing). Let representation be stochastic in any of the usual senses — (i) a randomized generator/recognizer; (ii) a probability assignment over expressions (e.g. a Solomonoff/algorithmic prior $M(x)=\sum_{p:\,U(p)=x}2^{-|p|}$); or (iii) things represented by probability distributions. Then still $|O_f| \le \aleph_0$, hence $O_n \neq \emptyset$ whenever $|G| > \aleph_0$.
Proof. (i) The strings a randomized machine can emit with positive probability, ranging over all seeds, form a subset of $\Sigma^{*}$, hence countable. (ii) A prior is a probability measure over the countable set of expressions, so its support — the positively-weighted representations — is at most countable. (iii) There are $2^{\aleph_0}$ distributions, but only the finitely specifiable (computable) ones are usable finite representations, and those are countable; an arbitrary real-parametrised distribution has no finite description and is itself a member of $O_n$, so option (iii) relocates irrepresentability without removing it. In each case the representable set injects into a countable set of finite specifications, so Lemma 2.1 applies unchanged. ∎
Let $E_h \subseteq G$, $E_h \neq \emptyset$, be the class of things experienced by humans. (We use non-strict $\subseteq$: properness is never invoked — see Q5.)
Theorem 4.1. If $|E_h| > \aleph_0$, then $E_h \cap O_n \neq \emptyset$ (indeed $|E_h \cap O_n| = |E_h|$). Equivalently, under Assumption A: if the space of human experience is uncountable, some experience is unverbalizable.
Proof. Suppose $E_h \cap O_n = \emptyset$. Then $E_h \subseteq O_f$. Compose the inclusion $E_h \hookrightarrow O_f$ with the choice $O_f \to \bigsqcup_i L_i,\ x \mapsto (i,w)$ for some $w$ with $\rho_i(w)=x$; this is an injection of $E_h$ into the at most countable set $\bigsqcup_i L_i$ (injective because $\rho_i$ is a function, so a codeword determines its referent). Hence $|E_h| \le \aleph_0$, contradicting $|E_h| > \aleph_0$. Therefore $E_h \cap O_n \neq \emptyset$. The refinement repeats the split of Prop. 3.1 with $E_h$ for $G$: $|E_h| \le |E_h\cap O_f| + |E_h\cap O_n| \le \aleph_0 + |E_h\cap O_n|$ forces $|E_h \cap O_n| = |E_h|$. ∎
And this is kind of where my thinking has ended. Going forward all of my work isRemark 5. Theorem 4.1 is conditional I guess. Everything then depends on:
\[ (\dagger)\qquad |E_h| > \aleph_0 . \]Remark 6 [compression form]. Read through Remark 2: $x$ is losslessly representable iff $\ell(x) < \infty$ iff $x \in O_f$. Theorem 4.1 then says: if $|E_h| > \aleph_0$, the coding map $c \colon E_h \to \{0,1\}^{*}$ cannot be total and injective — you cannot losslessly compress uncountably many experiences into finite codewords — so some $x \in E_h$ has $\ell(x) = \infty$: an incompressible experience, admitting no finite exact description. That is the ineffable element, and it is exactly the “compression function” the draft reached for, with the arrow corrected. The lossy counterpart sharpens the picture: language describes approximately, and rate–distortion theory gives, for a continuous source, $R(D) < \infty$ for every tolerance $D > 0$ but $R(D) \to \infty$ as $D \to 0$. One can always say something (finite bits, $D>0$) yet never say it exactly ($D=0$ needs infinite bits). A countable dense codebook approximates a separable $E_h$ to any $\delta > 0$; exact capture of uncountably many points is impossible. “Words approximate the sunset but never equal it” becomes the statement $R(0) = \infty$. (Full development in Q8.)
Remark 7 (set vs. proper class). If “all things” is a proper class, “$|G|$” is not a cardinal and Prop. 3.1 is ill-typed. The §0 fix presumes the totality can be set-bounded; it suffices throughout that $G$ merely contain an uncountable set (a copy of $\mathbb{R}$, read as magnitudes among “things”), to which the arguments localize.
Remark 8 (Richard, Berry, pointwise-definable models). Lemma 2.1 counts expressions under an externally fixed denotation, so $O_f$ is a genuine set of size $\le \aleph_0$ and the conclusion is immune to Richard's and Berry's paradoxes — which arise only when “is representable” is internalised as an object-language predicate (Tarski forbids this). One subtlety: there are models of ZFC in which every set is parameter-free definable (pointwise-definable models, Hamkins–Linetsky–Reitz); there the internal analogue of $O_f$ exhausts the universe and internal $O_n = \emptyset$. This does not contradict Prop. 3.1 (a statement about an externally fixed countable $\mathcal{R}$), but it shows “$O_n = \emptyset$?” is sensitive to the internal/external reading and to which $\mathcal{R}$ is admitted. Assumption A must fix $\mathcal{R}$ externally.
Q1. Does Gödel's contradiction criterion allow the creation of new information in the deterministic case? How would one prove it, question it, etc.?
?
Q2. Language follows rules, yet was created when there were no rules (or much simpler ones. How does that gel with Assumption A?
So read A extensionally (about what gets said): it survives, and powers the theorems. Read intensionally (about the saying), it is doubtful — and that doubt is not a problem but evidence for the program: the very capacity that creates language looks like an $O_n$ resident. Net recommendation: state A extensionally, and note that the generativity of language is itself a candidate ineffable.
Q7. “Things” vs. “information”: is there a better adaptation?
Your instinct is correct and worth keeping. In information theory “information” is formal by construction (it is bits/codewords/programs), so labelling the elements of $G$ “information” smuggles formalizability into the domain and makes $O_n = \emptyset$ by fiat — question begged. The clean fix is the Terminology note in §1: keep “things” for $G$ and define information $:= O_f$, the finitely codeable part. Then the thesis is the slogan things $\supsetneq$ information, and $O_n$ is “the non-informational.” This dissolves the worry: information is, correctly, the formal subset; the theorem says the world of things strictly exceeds it. Nearest existing vocabulary: Floridi's data → information hierarchy (data as primitive, pre-meaning differences) is the closest, but even “data” leans formal; “things”/“entities” stays maximally neutral. Keep “things,” and let “information” be the derived, bounded notion.