Monday, February 22, 2016

The Structure & Interpretation of Demonic Incantations

Some day, I shall get back to finishing up my thoughts on formal semantics and small languages; and some other day, I shall write more on programming language design. But today, I open up a new topic that I have not addressed on my blog before: creative writing! More specifically, hard-fantasy worldbuilding.

I like magic with rules. I like fantasy stories in which magic isn't "magic"- it's just the way the world works, part of physics that just happens to be different from ours. As a programmer myself, I'm also fond of fantasy stories with computers in them- things like The Silicon Mage and parts of the Castle Perilous books. The astute reader may recognize the title as a play on Structure and Interpretation of Computer Programs. And I'm certainly not the first to invent a magic system based on computer programming, but I think I've got a pretty fun personal take on the concept.

The Background

Several years ago, Eliezer Yudkowsky published a short story called Initiation Ceremony, a followup to the essay "To Spread Science, Keep it Secret"; the central idea here is that people like knowing secrets and being part of conspiracies- so if we want to increase scientific literacy and respect for science, maybe the best route would actually be to keep it secret! Or, put on a show of doing so, at least.

That got me thinking- what if there was a real practical reason why some particular science had to be kept secret?

Maybe the universe itself is programmable, and that fact needs to be hidden from malicious hackers (not terribly unique yet).

Maybe the universe is programmable, and it's ridiculously easy to do, if you know the trick. So easy that literally anybody can access the super-user interface with 5 minutes instruction to tell them how; but if you don't know what you're doing, an improperly educated person could accidentally destroy the world....

That's something worth keeping as a very carefully controlled secret. Plausibly, that would be very difficult to do indefinitely. So perhaps we're looking at a world that has undergone multiple cycles of civilization and destruction, whenever the secret gets into the wrong hands.

So, how does the programming interface work? Why, by giving instructions to demons of course! Humans acquiring magical powers by making deals with supernatural creatures- demons, angels, or what-have-you- is an old and respectable fantasy trope; we just need to set up some rules that make it more like computer programming, and less like Dr. Faustus. If demons are really good at following directions, but natively really stupid, like computers....

In more detail: Presume that the universe is animist, down to the very lowest levels. Basically, it works like our world, but the thing that "breathes fire into the equations and makes a universe for them to describe" is that every particle is actually an animist intelligence which has been instructed in the rules of physics that it is meant to follow (by whom!?). Additional intelligences exercise hierarchical control over various higher levels of organization- most notably, there's a primordial intelligence, or spirit, behind every living organism, including every human, and a lot of inanimate objects too, a fact which would be objectively, scientifically verifiable to a properly-trained thaumatologist. If you could talk to the intelligence behind an electron, or anything else, and tell it to behave differently, it would, just like you can talk to a person and convince them to do things- but how do you talk to an electron? You need some kind of intermediary, someone or something that can communicate directly with the fundamental intelligences of the animist world, but which can also understand you, the magician.

Now we get into some theology- an empirical science in this world, even if not everyone believes in precisely the same religion! Clearly, someone had to give the electrons and such their initial instructions; the world was formed out of a chaos of primordial intelligences at all levels of advancement by a Creator, who picked the simplest intelligences for the jobs of being fundamental particles, and also instructed them to obey certain kinds of commands under certain circumstances from more advanced intelligences, thus providing a mechanism for the intelligences selected to serve as human spirits (and those of other animals) to exercise free will over their bodies. This implies two things: first, that even really simple primordial intelligences must be aware of each other and capable of communicating with each other, and also capable of solving fairly complicated mathematical problems extremely rapidly; and second, that, given the finite number of animals in the world, that there is still an infinite number of advanced intelligences left in the primordial chaos outside the world.

Some of that infinite sea of left-out intelligences would probably like very much to have the chance to interact with Creation.

Since there is a Creator who told the world how to work in the first place, miracles are easily performed if that Creator just decides to give some updated instructions every once in a while; or, alternatively to instruct the animist world to understand and obey commands from certain delegated servants- priests of the one true religion. That's one kind of "magic" that would exist in this world, but so far it just looks like a monotheistic religion in a world indistinguishable from our own. In order to turn "magic" into a reproducible, objective science, we presume that, having set up humans as lords over his Creation, this Creator has also given every human the authority to allow additional intelligences from the infinite chaos permission to interact with the world, and give them delegated authority over some material domain to serve as their "body".

The only requirements are that a human
a) be "sane" (another fiddly concept which would suddenly have an objective, verifiable test in this world!)
b) genuinely want such an intelligence to enter the world.
c) genuinely believe that it will work

That gets you the most basic, simplest possible "magic spell"- the equivalent of "uum" from Sam Hughes's serial novel Ra, a spell that "works" but does nothing, as no instructions are provided to the summoned intelligence.

Once you add the intention that the summoned intelligence have permission to interact with the material world in any way, however, that's when things get dangerous! Before you meet them (and remember, they exist in the infinite chaos outside the world- where are you going to get the chance?), there's no way to select a specific intelligence to summon; whichever random one you get might be fairly smart, or insufferably stupid; and it might be naturally benevolent, or excessively malevolent, or simply ignorant- a state often indistinguishable for malevolence! If your instructions are either impossible to carry out, or insufficiently precise, you might get lucky with having summoned an "angel"- a benevolent, reasonably smart intelligence- but much more likely, the results will be some level of disastrous. Hence, the common name for summoned intelligences: Demons. It's just safer to assume that they're all dangerous and malevolent, and make sure you know how to control them properly.

The requirements for doing magic- believe that it will work, and wish for a demon to show up with a particular set of instructions- are so simple that they could be easily rediscovered by just about anyone with a sufficiently open mind. "Fortunately", much like the sparks from Girl Genius, most people who independently rediscover demon summoning without proper training are likely to accidentally kill themselves, or at least get scared off it, well before they can do any large-scale damage. But, to really maintain a stable society would require a massive over-arching conspiracy and disinformation campaign about the true nature of magic, such that people do not actually believe that they are capable of it, rather like the situation in Shin Sekai Yori, where all humans possess powerful telekinesis.

And, if initiates to the conspiracy want to be able to actually use magic at all safely, it would be necessary for them to develop something very like our computer science, in order to ensure that they can provide absolutely precise instructions to obviate any possible differences in the personality or intelligence level of any summoned demon, and ensure that they will all behave precisely the same, predictably and reproducibly- and safely- for any given spell.

Additional Restrictions

If demons can be used to re-instruct any matter arbitrarily, as the Creator of the World could do, that's a little over-powered- not to mention difficult for humans to practically figure out how to use. Magic systems are made interesting not so much by what they can do, but what they can't do- and how characters manage to exercise their creativity to get around those restrictions anyway.

First, the low-level limitations: demons can produce effects equivalent to applying arbitrary forces, or manipulating statistical outcomes, but they are not permitted to violate certain fundamental invariants of the universe, such as:
  1. No effects that would violate conservation of energy or momentum.
  2. No transmission of information faster than light.
  3. No exact cloning of quantum information.
etc. That last one means you better not ask a demon to "make an exact copy" of something, even given the necessary raw materials to work with. After all, an exact copy would require duplicating a quantum state. Duplication spells are going to be a little more complicated than that! These rules would be enforced by the original instructions given by the Creator to the animist intelligences that form the physical world. Understanding and command of magic in this world will, therefore, go hand-in-hand with advancing knowledge of physics to explain what exactly demons can do, and why they can't do certain things.

Additionally, the source of magic (essentially, delegation of creative power to humans by divine fiat) suggests a few other higher-level rules:
  1. A demon cannot directly alter any part of a human's body without explicit consent of that human.
  2. A demon acting on its own cannot intentionally kill a human.
But, that doesn't mean a demon can't be instructed to do something else that the magician knows will end up killing someone, nor that a demon taking malicious or ignorant advantage of unclear instructions can't make life incredibly unpleasant, and perhaps accidentally fatal, for the summoner and those nearby. Perhaps they simply find it amusing to replace all of the air in the room with jello.... But they do prevent accidental world-wide catastrophes due to malicious demons summoned by untrained magicians. These rules would be actively enforced by some of the higher-level controlling intelligences of the world- "guardian angels", if you wish.

A Specific Magical Society

Under this framework, there are still all kinds of different societies that could develop. That's part of what makes for an interesting world, in which all kinds of different stories could be told about different people in different places. Here's one that I find interesting:

As previously noted, this world has probably gone through several catastrophes, where some rogue magician has had the power to "destroy the world". The last disaster left behind a slew of magical artifacts from the previous civilization- things which, on a purely physical level, are unremarkable, but which have demons attached to them running complicated "software" and just waiting around for someone to activate the programmed user interface appropriately. These provide for a rich source of ad-hoc folk magic beliefs, given that they don't behave according to any necessary logical rules- merely according to the whim of the ancient magician who created them, and whatever they thought "made sense". They would also be intriguing objects of study for specialists in reverse-engineering, attempting to use the ancient objects to advance the state of modern thaumatology.

The conspiracy and practice of magic is controlled by a Guild of Magicians, who do not publicly reveal the source of magical power, but are very clear on the fact that it's extremely difficult to explain or make use of, certainly not something that just anybody can do, although they are happy to welcome anyone willing to go through the necessary many years of preliminary training. In order to create a supply of adepts who are appropriately trained in the mental discipline of composing precise and unambiguous instructions, they accredit and control Colleges of Thaumaturgy in universities through the country, one of whose departments is of course Algorithmics- of which the basic reference work is, of course, The Structure and Interpretation of Demonic Incantations. Demons, of course, are purely theoretical constructs- oracles that one may assume have no real intelligence but will follow any instructions given them quickly and precisely, and whose operation can be approximated by certain useful algorithmic devices studied in the College of Engineering. Unlike most CS programs in our world, however, courses in Algorithmics would focus heavily on formal provability of incantation (program) correctness and security.

An accredited degree in algorithmics, plus some study of physics, qualify a student to apply for entrance to the Guild, to assist a practical magician. Those who prove sufficiently skilled and trustworthy are eventually advanced to the rank of Licensed Magician via a private initiation ceremony in which the secret is revealed- that demons are real, and your incantations will become effective (i.e., your programs will run) if you believe it, and wish it. After that point, they may begin to work as practical magicians in some field of engineering, or as academic researchers.

Occasionally, of course- every couple of years or so- someone rediscovers the secret anew anyway, or some disgruntled magician really does leak it to their friends despite their extensive training, and a little bit of havoc ensues. When that happens, the Guild must spring into action to effect containment of the disaster, and construct a cover story about how the secret of magic was leaked, and reiterating to the public just how dangerous it is outside of the hands of the Guild, and only the Guild.

Of course, there are other countries in the world, not all of which are under control of the Guild of Magicians, but that's no worry- all of the ones that still exist clearly have their own effective means of keeping the secret under control, and if they don't tip off their own citizens, they certainly won't tip off ours....

Tuesday, September 22, 2015

A Progressive Model of WSL Syntax & Interpretation: Part 3

Last time we saw how to introduce generalized quantifiers and arbitrary specifier positions into our model of WSL; but, in the process, we lost any recognition of the scoping effects between quantifiers. Figuring out how scoping effects work for generalized quantifiers represented by sets-of-sets can get pretty complicated and confusing, so we're gonna really slow and step-by-step.

First, let's revise and review our model of the syntax so far. The complete syntactic model that I'll be using for the rest of this post is given by the following grammar:

S  → P AP
AP → RP AP | 0
RP → QP R | QP e
QP → Q NP
NP → N NP | 0

Where an S is a Sentence, an AP is a Argument Phrase, an RP is a Role Phrase, an R is a Role, a QP is a Quantifier Phrase, a Q is a Quantifier, an NP is a Noun Phrase, and N is a Noun.

Next, let's look at some examples to get a better grasp on how scope should work. Consider the English sentence

"Everybody loves somebody."

This has several possible translations into WSL; two of them are:

1) Ka ves anz i siru jest anz jo.

and

2) Ka jest anz jo siru ves anz i.

Sentence (1) states that, for every person, there is someone whom that person loves- but not every lover necessarily loves the same lov-ee. Sentence (2), on the other hand, states that there is some single person whom everybody else loves, because the existentially-quantified patient is now outside the scope of the universally quantified agent.

Additionally, sentence (1) allows that every agent might be participating in a totally separate instance of "loving", while sentence (2) indicates that there is only one instance of "loving" going on, and every person is a simultaneous agent in it. This is because of differences in the scope of the existentially-quantified specifier ("siru", with a phonologically-null quantifier). Moving "siru" to different positions produces more subtly-different interpretations; e.g., taking (2) and moving the specifier to the end would indicate that there is one person who is separately and independently loved by everyone- possibly at different times. And adding multiple conjoined specifiers would make things even more complicated.

Now let's take a look at how Roles get assigned to QPs inside Role Phrases, as of last time:

[|RP: QP R|] = λx. {y : ∃z. z ∈ x & [|R|](z)(y)} ∈ [|QP|]

What we're doing here is constructing the set of all things that bear a particular relation to some element of the specifier set, and then asserting that that is the same as one of the potential referent sets from the Quantifier Phrase. This implicitly imposes some constraints on the identity of the specifier set as well. A Role Phrase containing a specifier works a little differently:

[|RP: QP e|] = λx.∃Y. Y ∈ [|QP|] & Y ⊆ x

This just imposes an explicit constraint on the identity of the specifier set. (Note that I have chosen here to use an upper-case Y for the referent set in the rule for specifiers, to distinguish it from the lowercase y used for an element of the referent set for normal Role Phrases.)

In order to respect quantifier scoping, we need to arrange things so that we can reconstruct all lower-scope quantifier sets and select the relevant constraints independently for every element y of the referent set that we're constructing for the current RP.

In order to do that, we first have to make the denotations of all lower-scoped RPs available during the interpretation of any given RP. That means modifying our interpretation rules for APs as follows:

[|AP|] = λx.[|RP|](x)([|AP|])

such that the denotation of next-lower-scope Argument Phrase (which contains all the remaining Role Phrases) is filtered through the current Role Phrase as a parameter. That will, of course, require updating the rules for RPs to take multiple arguments, and doing the right thing with them:

[|RP: QP R|] =
        λx.λa. {y : ∃z. z ∈ x & [|R|](z)(y) & a(x)} ∈ [|QP|]

[|RP: QP e|] =
        λx.λa. ∃Y. Y ∈ [|QP|] & Y ⊆ x & a(x)


This places the evaluation of each Argument Phrase inside the scope of the variable y bound by the next higher Argument Phrase. Although y itself is not accessible in the lower scopes (we only pass along the shared specifier set as a parameter to a), this means that lower quantifiers are re-evaluated for every y. Thus, in "Ka ves anz i siru jest anz jo.", it is possible in the evaluation of "jest anz jo" to select a different existentially-quantified referent to correspond to every member of the universally-quantified referent set of "ves anz i". This is obvious in the semantics for specifiers, where we're still cheating a little bit by using the "∃" symbol to bind the variable Y and borrowing its scoping behavior.

The interpretation for the internal structure of a QP, containing a Q and an NP, remain unchanged. If we round things out with an explicit rule for null APs, we get the following completed model for the syntax-semantics interface:

[|S|]  = [|P|]([|AP|])

[|AP: RP AP|] = λx.[|RP|](x)([|AP|])
[|AP: 0|]     = λx. true

[|RP: QP R|] =
        λx.λa. {y : ∃z. z ∈ x & [|R|](z)(y) & a(x)} ∈ [|QP|]

[|RP: QP e|] =
        λx.λa. ∃Y. Y ∈ [|QP|] & Y ⊆ x & a(x)


[|QP|] = [|Q|]([|NP|])

[|NP: N NP|] = [|N|] ∩ [|NP|]
[|NP: 0|]    = U

[|P|] = G[P]
[|R|] = G[R]
[|Q|] = G[Q]
[|N|] = G[N]

This gets us the ability to model a pretty big chunk of all WSL declarative sentences. Still to come: non-intersective Nouns, modality, alternative Projectors, subordinate clauses, and controlling semantic projection.

Monday, September 21, 2015

A Compositional "yes" and Other Discoveries

Since WSL is radically everything-drop, it recently occurred to me that a complementizer (or "projector" in WSL-specific terminology) and a modal particle with no other content words nevertheless constitute a complete sentence (in fact, just a bare complementizer should constitute a complete sentence, but I decided that would be pragmatically weird in every case, and just Not Done).

So, of course, I had to set out to figure out what all the combinations would actually mean. Not as an exercise in formal semantics, but idiomatically; how would native speakers of WSL actually use these short phrases in everyday life?

The four basic modal particles are

es - normal realis mood
miy - approximately equivalent to "could be" or "might, for all I know".
bi - logical possibility
pek - roughly equivalent to "should" or "it better be"

There are ten projectors, so that gives a total of 40 possible combinations. 16 of them (plus two more we'll consider later) can be complete sentences on their own:

k'es (ka+es): an assertion that some contextually-provided proposition is true. I.e., "yes". But not all the time! We'll come back to this later....
ka miy: "Sure, if you say so."
ka bi: "That is definitely a logical possibility." (I expect this one is typically used sarcastically, with the implication of  "no, I don't really think so".
k'pek (ka+pek): "If not, we got problems."

tc'es: "No". But again, not all the time....
tce miy: "That is probably wrong"
tce bi: "Not necessarily"
tce pek: "I hope not"

em es: "Is it?"
e'miy: "Could it (for all you know)?"
em bi: "Could it ever?"
em pek: "Should it?"

mi's: "Ain't it?"
mi miy: "Couldn't it?"
etc, You can guess the last two.

While figuring those out, I also came up with this bonus phrase, which includes an extra third word:

ka ves bi: "This is an obvious logical necessity in all possible worlds".
Or, more colloquially: "Well, duh!"
(This sentence, small as it is, is actually structurally ambiguous, but the alternate reading
is the very weird "it is a logical possibility that everything is that thing". Not something you'd need to say very often!)

The next 16 combinations are not complete sentences, but are valid nominal clauses:

vr'es (vor+es): Something. Anything. I'm thinking this might end up getting used as a generic indefinite pronoun for when you really don't care what (cf. "mek", which is the indefinite pronoun "one", but frequently gets used as a  cataphor for dislocated nominal clauses)
vor miy: a hypothetical entity which you posit to exist, but might not
vor bi: any hypothetical entity, whose actual existence is irrelevant
vor pek: that which darn well ought to be. Like flying cars.
Em intc mot "flying cars" jesihu!? K'ajnu vor pek es!
Where are my flying cars!? We should have those by now!

votc es: a nonexistent thing. I'm thinking this might be an idiom comparable to "unicorn"
votc miy: a thing which you are really darn sure does not exist. The emphatic "rainbow-vomiting nuclear unicorn", for example.
votc bi: a logical impossibility
votc pek: that which darn well oughtn't to be. Like social spiders... ick!
Ka votc peku "social spiders" es!
Social spiders are a thing which should not be!

vm'es (vem+es): "whether it is"
ve'miy (vem+miy): "whether it could be, as far as you know"
vem bi: "whether it's logically possible"
vem pek: "whether it should be"

vmi's (vmi+es): "whether it isn't"
vmi'y (vmi+miy): "whether it isn't, as far as you know"
etc.

Finally, the following combinations can be either complete stand-alone sentences, or nominal clauses:

s'es (sa+es): "Yes, they are" or "the fact that they are", asserting that a predicative relationship holds.
satc es: "No, they aren't" or "the fact that they aren't"

You can probably fill in the remaining 6 combinations for the other moods.

So, not only have I discovered the WSL words for "yes" and "no, we've also found that there are two different ways of saying "yes" (k'es and s'es) and two different ways of saying "no" (tc'es and satc es), conditioned on whether the question you are responding to is about a predicative relationship or not.

Additionally, the internal structure of a compositional "yes" or "no" parallels the syntactic structure of the question. So, if somebody asks you a negative question, like "Aren't you going?" responding with "Tc'es" means "No, I'm not"- confirming that the negative that the questioner used was
appropriate. If, however, you respond with "K'es", that actually means "Yes, I am"- contradicting the questioner's use of a negative. Thus, there is no confusion over "was that a 'yes, I'm not', or a
yes-like-'no, I am'?", and no need for yet another word like the French "si" just to take care of answering negative question unambiguously.

A Progressive Model of WSL Syntax & Interpretation: Part 2

Last time, I ended with the note that properly modelling specifier phrases would require splitting the interpretation of argument phrases in half; in particular, I had in mind the idea that variable bindings for noun phrases would need to be moved around to ensure that the specifier variable would be in-scope in the semantics for every argument phrase.

It turns out that the solution is actually much simpler. First, we will introduce a very simple change to the syntax rule for a sentence to account for sentential Projectors (a part of speech which heads independent clauses in WSL):

S  → P QP e AP

Next, we'll stop treating the specifier phrase separately, and account for it as a special case of an argument phrase:

S  → P AP
AP → A AP | 0
A  → QP R | QP e

We could also choose to treat the specifier clitic as a kind of Role, which would be slightly simpler, but this formulation better reflects my own psychological perception of what a specifier is (and thus presumably reflects the intuition of the fictional native speakers of WSL as well). Note that this allows a single clause to contain multiple specifiers, as well as putting them in arbitrary positions with respect to the other arguments; that situation is accounted for in the WSL Primer, which says that the semantics of multiple specifier phrases is the same as that of multiple conjoined specifiers, except that using multiple specifiers allows you to place them all in different quantifier scoping levels- the same as the interpretation for repeated roles.

The semantics for these bits of syntax is as follows:

[|S|]  = [|P|]([|AP|])
[|AP|] = λx.[|A|](x) & [|AP|](x)
[|A: QP R|] = λx.[|QP|](λy. [|R|](x)(y))
[|A: QP e|] = λx.[|QP|](λy. y ⊆ x)

To summarize: the denotation of a sentence is the denotation of the argument phrase chain filtered through the denotation of the Projector; the denotation of an argument phrase is the denotation of the argument given an entity variable x conjoined with the denotation of the remaining argument phrase given x; and the denotation of an argument is the denotation of a relation on the shared entity variable x and the phrase-specifier entity variable z filtered through the denotation of the quantifier phrase (which for the moment is unchanged from last time).

The case of an argument containing a QP and specifier clitic instead of a QP and a Role just contains the explicit relation that the phrasal entity z is a subset of x. Note that this requires interpreting z and x not as representing single referents, but as sets of possible referents- an idea I introduced earlier in the series on semantics for a monocategorial language.

For now, we'll only deal with a single Projector: "ka", which indicates a simple declarative sentence. It's semantics are very simple:

[|P|] = G["ka"] = λy.∃x. y(x)

This just says that some x (which now refers to a set) exists, and its identity will be constrained by y (which is bound to the denotation of the argument phrase chain). We could build this into the interpretation of an S directly, but we will have to deal with the semantics of Projectors at some point, so we might as well start here.

We now have the ability to intersperse specifiers and other arguments in any order, with the members of the specifier set constrained by the specifier phrases, and the scope of all quantifiers corresponding exactly to their surface order!


Generalizing Quantifiers

The next step in building our model of WSL semantics is to improve the handling of quantifiers. As described in this article, the denotation of a Quantifier phrase will be represented not by a proposition in predicate logic, but by a set of sets of possible referents- all those sets of referents which contain the appropriate quantity of the type of referent identified by a given Noun phrase. This helps our model conform to the intuition that a bare noun or quantifier phrase does not correspond to a logical assertion- i.e., that a given entity exists- but merely to a possible entity itself. The semantics for any individual quantifier are given by a function which takes in the denotation of a Noun phrase and uses it to construct the appropriate set of sets. The new interpretation rule for QPs is as follows:

[|QP|] = G[Q]([|N|])

And some examples of Quantifier semantics are as follows:

G["ves"] = λy. {x ⊆ U : y ⊆ x} ("every" or "all"; i.e, the set of all sets in the universe U that contain the entire set y)

G["jest"] = λy. {x ⊆ U : |y ∩ x| > 0} ("some"; i.e, the set of all sets in the universe that contain at least on element of y)

G["hiq"] = λy. {x ⊆ U : |y ∩ x| = 5} ("five"; i.e, the set of all sets in the universe that contain exactly some five elements of the set y)

This of course also requires a new formulation of the semantics for Noun phrases that produces the basis set of "referents of the right type". The new Noun rule is as follows:

[|N: n N|] = G[n] ∩ [|N|]
[|N: 0|]   = U

Rather than binding a new entity variable which we assert to satisfy a given predicate, or to be a member of a given set (given by looking up the Noun n in the lexicon), we simply directly construct the intersection of all of the sets that are the denotations of individual Nouns.
Finally, we have to update the interpretation rules for arguments (again) to handle the new kind of denotation for QPs:

[|A: QP R|] = λx. {y : ∃z. z ∈ x & [|R|](z)(y)} ∈ [|QP|]
[|A: QP e|] = λx.∃y. y ∈ [|QP|] & y ⊆ x

This says that a normal argument containing a role marker asserts that the set of all entities y which satisfy a particular relation with some element of the specifier set x is in the denotation of the QP; or, that an argument containing the specifier clitic asserts that some element y of the the denotation of the QP is a subset of the specifier set x.

Unfortunately, we've just undone our progress in allowing for the correct quantifier scopes! I constructing a compositional semantics for quantifiers in terms of sets, we have thrown out the conventions for establishing variable scopes in predicate logic- because we eliminated the variables! In order to recover the proper quantifier scopes, we're going to have to find a way to take into account the constraints imposed in lower scopes while constructing the sets for higher-scoped quantifiers.

We'll take a look at that problem in the next post in the series.

Sunday, September 13, 2015

Non-intersective Nouns & Negative Scopes

This post is a follow-up to several prior discussion on formal semantics for a minimalistic monocategorial language. If you aren't familiar with them already, you may want to read part one and part two before proceeding.

Non-intersective Nouns

There is a class of adjectives called "non-intersective adjectives" because the set of referents of a noun phrase that includes them does not intersect the set of referents for the same noun phrase without.
This concept is best explained with examples: "a short basketball player" is also "a basketball player", so "short" is an intersective adjective--the set of "short things" intersects the set of "basketball players", and the meaning of the whole phrase is the intersection of those two sets.
On the other hand, "a former basketball player" is not "a basketball player", so "former" is non-intersective. The meaning of "former basketball player" is not a subset of "basketball player", and isn't formed by intersecting it with anything. It's a completely disjoint set, but one that does have a logical relationship to the set of "basketball players".

Similarly, if you "almost finish", then you do not "finish", so "almost" is a non-intersective adverb.

My formal education in formal semantics only really exposed me to one way of modelling non-intersective adjectives/adverbs: as higher-order functions that take predicates as arguments and produce new, different predicates. Thus, the predicate logic notation for a "former player" is former(player')(x) (vs. player(x)) , for some entity x.

But this model is not very compositional (i.e., if a "former player" is former(player')(x), a "former basketball player" is... what exactly?), and results in icky complicated interpretation rules. As a result, I have so far mostly avoided them in WSL, and the few times I've needed them I've punted
and just decided that they are taken care of by morphological affixes. That works as long as you just want to say something like "a former player", and leave the "basketball" (or any other additional descriptors) out of it.

Contemplating the programming language Prolog, however, provides a way out of this mess. Basic versions of Prolog do not have functions that can compute arbitrary values--just predicates that can be true or false. A Prolog system can, however, simulate functions by allowing you to query it about what the possible values are that would need to be plugged in to one or more argument positions of any given predicate in order to make it evaluate to "true". You can specify "input" values for whatever arguments you want, and for whatever arguments you leave unspecified, the Prolog system will give you a set of all possible values of those variables that satisfy the predicate as "output". Since Prolog is based on first-order logic, the exact same transformation for simulating functions works for modelling non-intersective adjectives in a predicate logic model for formal semantics.

So, non-intersective adjectives can be modeled as two-place predicates (with one input argument and one output argument) that specify the relation between the actual final referent of a noun phrase and the thing-that-it-isn't. Plus some mathematical machinery to glue it all together. That insight will allow us to add the capacity for translating non-intersective adjectives into our monocategorial language. I say "translate" because, being monocategorial, the language does not have a distinct class of "adjectives" to add non-intersective members to. Instead, it will have "non-intersective nouns" (or "non-intersective noun-jectives").

Introducing that machinery into our monocategorial language requires a little bit of updating to the existing interpretation rules. The altered syntactic rules are as follows:

[|P: w P|] = λx.λy.λz.[|w|](x, y, z, [|P|])
[|P: w ,|] = λx.λy.λz. [|w|](x, y, z, λx.λy.λz. y ⊆ z & ∃r. y ⊆ {z : ∃e. e ∈ x & r(e, z)})

Essentially, rather than immediately evaluating the denotation of the current word and the rest of the phrase with the same arguments, we pass the denotation of the rest of the phrase into the semantic function for the current word, which will allow the lexical semantics of a word to control what arguments are given to the rest of the phrase. At the end of the phrase, we pass in the "default" expressions that account for the possibility that no explicit quantifier or relation words were present. Note also that I have introduced the comma-separated argument list notation as "sugar" for repeated application of a curried function, since the large number of parameters to our lexical semantic level has started to become unwieldy.

Of course, since we've changed the lexical semantic interface in the syntactic interpretation rules, we have to also update the templates for our lexical semantic classes. The existing classes are updated as follows:

a) λx.λy.λz.λp. z ⊆ red & p(x, y, z)
b) λx.λy.λz.λp. y ⊆ {w : ∃e. e ∈ x & ag(e, w)} & p(x, y, z)
c) λx.λy.λz.λp. x ⊆ run & p(x, y, z)
d) λx.λy.λz.λp. x = y & p(x, y, z)
e) λx.λy.λz.λp. |y| > |z - y| & p(x, y, z)

In every case, we simply pass along all the original arguments directly to the semantic function for the remainder of the phrase, and logically conjoin it to the original lexical semantic expression. However, we now have the option of messing with those arguments if we so please; thus, we can introduce an additional semantic class for non-intersectives:

f) λx.λy.λz.λp.
        y ⊆ z & ∃r. y ⊆ {z : ∃e. e ∈ x & r(e, z)} &
        ∃x'.∃y'.∃z'. z ⊆ {a : ∃b. b & ∈ y' & former(a, b)} &
        p(x', y', z')

There's a lot going on here, but most of it is just book-keeping boilerplate. Let's break it down:
First, we cleanly terminate the description of the current referent; by the time we get to the end of the phrase, we won't be talking about the same thing anymore, so we have to throw in all the "just-in-case" assertions (that the referent set is some subset of the quantifier base and that it is has some relation to the sentential event) in here as well. In the next line, we assert the existence of a new quantifier base and a new referent set (z' and y', respectively) and, crucially, a new event x'--because if we're talking about a "former basketball player", then the "basketball player" which-he-isn't doesn't have any relation to this clause's event, so we have to replace it with something else. We then assert that all members of the current quantifier base have a specific relationship (in this case, "former") to some member of the new referent set. And finally, we use the rest of the phrase, denoted by p, to describe the new referent and new event.

The addition of this mechanism to the language has a few interesting long-range consequences. The most obvious is that word order actually matters. Up until now, we could've cheated on the phrase-structure rule and just said P → w w* , using the Kleene star operator for arbitrary repetition, instead of P → w P | w ,; but now, the fact that later words are in smaller phrase embedded at a lower syntactic level than previous words is very significant. Changing the ordering of words around a non-intersective noun can change which referent those words are actually describing. Similar effects are present in English; a "former blue car", for example, is not necessarily the same thing as a "blue former car". The first one may still be a car, but of a different color, while the second is definitely no longer a car, but definitely is blue.

Note that while syntax now encodes new information in word order about referent scope, the syntactic rules no longer encode information about logical connectives; the implicit conjunction of all lexical items is now a function of lexical semantics instead. We could consider this an accidental artifact of the formal framework we're using, but we might also come back to it later and exploit the lexification of logical connectives to come up with some new lexical semantic classes.

The second long-range consequence is that class-c words now have a real essential function; in the beginning of a phrase, prior to any non-intersectives, they are essentially adverbs, unconnected to the phrasal referent, which can float between phrases with no change in sentential meaning. After an intersective, however, they no longer act on the clausal event. Instead, they describe some event that the new referent is a participant in. The same applies to class-b and -d relation words.

Negative Scopes

Most non-intersectives have an implicit "not" built in to them. A "former basketball player" is not a "basketball player", a "fake Picasso" is not a "Picasso", and while an "alleged thief" might be a "thief", he also might not. And it turns out that with the interpretive machinery we've built so far, we can actually translate "not" as a non-intersective noun with the following lexical semantics:

not: λx.λy.λz.λp.
        y ⊆ z & ∃r. y ⊆ {z : ∃e. e ∈ x & r(e, z)} &
        ∃x'.∃y'.∃z'. z ∪ y' = {} &
        p(x', y', z')

I.e., "the relation between the quantifier base for the current referent set and the new referent set is that they have no elements in common".[1]

We can thus expect all non-intersectives to behave somehow similarly to negatives in the behavior they trigger in lower scopes.[2] This gives us a possible use for semantically-significant stress focus, in addition to the rising-falling intonation patterns that are used to denote phrase and sentence boundaries. We can use stress focus to indicate the specific reason that a referent "is not" something else. Referring to a previous example, the ambiguity of "former blue car" can be resolved by stating that it is either a "former blue car" or a "former blue car"--indicating that the reason the item in question is "former" is because it is no longer "blue" in the first case and no longer a "car" in the second.

More traditionally, we might want to make some special accommodations for phrases that are described by monotone-decreasing quantifiers, like "no" or "few", which are also "negative", in a slightly different way. Perhaps the "reason" for a decreasing quantifier can also be indicated by stress focus (do "few men work" or do "few men work"?) The practical impact of that level of ambiguity ("does this instance of focus refer to the quantifier or the non-intersective scope?") is likely to be minimal.

In either case, this will already be a very long blog post, so updating the syntax to recognize focused items is left as an exercise for the reader.

The Weird Ones

Non-intersective nouns, it turns out, don't have to be used only to translate what English encodes as non-intersective adjectives. They're just words that specify some arbitrary not-necessarily-intersecting relationship between the quantifier base for one referent set, and some other referent set. Y'know what else acts like that? Adnominal adpositions. E.g., prepositions that describe noun phrases. Also, genitives (for which English can use the preposition "of", but doesn't always).

Wanna say "Bob's cat climbs trees" in monocategorial form? "Cat agent of Bob, tree theme, climb event." The word "of" ends up as a non-intersective noun! Crazy! Incidentally, its semantics are as follows:

of: λx.λy.λz.λp.
        y ⊆ z & ∃r. y ⊆ {z : ∃e. e ∈ x & r(e, z)} &
        ∃x'.∃y'.∃z'. z ⊆ {a : ∃b. b & ∈ y' & ∃r. r(a, b)}
        p(x', y', z')

I.e., "all members of this quantifier base have some kind of relation with some member of the new referent set".

And there are all sorts of normal English nouns that entail a relationship to some other unspecified referent. "Father", for example. A "father" cannot exist without a child, so the concept of "father is naturally expressed in terms of a two-place predicate... which will be embedded inside the machinery for non-intersectives to allow describing both halves of the relationship, should you so desire[3]. So, "Bob's father build's houses"? In monocategorial form, it's "Agent father Bob, many house patient, build event". Note that in this case, the word order in the phrase "Agent father Bob" is critical. If we swap things around, we get the following different readings:

"father agent Bob": "Bob-who-does-something's father is somehow involved"
"Bob father agent": "Bob, who is the father of someone who does something, is somehow involved"
"Bob agent father": "Bob, who is a father, is the agent (builds houses)"

And if we want to cover both sides of the relationship at the same time: "John agent father Bob, many house patient, build event" ("Bob's father, John, builds houses.")

[1] Note that this is not the same as a translation for "no", the negative quantifier, which is addressed a little further down the page. This sort of "not" means "I'm talking a referent that is identified by not being that other thing", as opposed to saying "no referents of this description are involved in the action".
[2] While similar, this is actually not the same thing as the "negative scopes" that license negative polarity items like "anymore" in English. Those are a feature of the behavior of certain quantifiers, like the negative quantifier "no" described in [1].
[3] WSL encodes these kinds of nouns as Roles. When we get back to the semantic model for WSL in a later post, we'll see the utility of a productive morphological derivation system to turn Roles into non-intersective Nouns and vice-versa.

Tuesday, September 8, 2015

Generalized Quantifiers for a Monocategorial Language

At the end of yesterday's post, I briefly mentioned the concept of generalized quantifiers. Today, I want to investigate how the semantics of the minimalistic language can be extended with that concept to eliminate the need for "built-in" existential quantification.

The first step is to recognize that noun phrases do not necessarily denote single referents. In a sentence such as "Every student goes to school", for example, the phrase "every student", while grammatically singular, does not refer to just one student; rather, it denotes a set of students (in this case, all of them), and the sentence makes a statement about the properties of that set- all of it's members also belong to the set of things that go to school. Or, in other words, [|every student|] is a subset of [|goes to school|].

We can thus modify yesterday's semantics so that all entity variables actually refer to sets. In this case, the form of the syntactic interpretation rules need not change (although how we read them may be tweaked, replacing "There exists some x" with "There exists some set x"), but lexical semantics for each possible word type end up looking like this:

a) λx.λy. y ⊆ red
b) λx.λy. y ⊆ {z : ∃e. e ∈ x & ag(e, z)}
c) λx.λy. x ⊆ run
d) λx.λy. x = y

Note that predicates are defined by the set of arguments for which they are true. Thus, predicates are set-valued entities on which we can use set operators like ⊆ (subset) and ∈ (element of); the traditional predicate logic notation that we have been using so far, pred(x), is simply shorthand for the set-theoretic formula x ∈ pred.

The only major alteration introduced here is seen in the semantics for relation words, class b, which must be modified to explicitly construct the subset of entities which have a particular relation to some element of the set of events.

It is now straightforward to add a fifth semantic class of words which in some way restrict the cardinality of (or quantify) a set:

e) λx.λy. |y| = 5

(There could be an additional sixth class that restricts the cardinality of the event set, represented by the variable x, but for simplicity we will ignore that possibility for now.)

This works great for simple numerals (like 5, as shown in the example) and more vague things like "many" or "a few"; but, it causes problems for quantifiers like "most" or "one-third" which restrict the cardinality of the referent set compared to what it would have been if it were not quantified. We need some way of keeping track of that original maximal set.

In a more "normal" language, generalized quantifiers would operate at a separate syntactic level from nouns, and could take in the compositional denotation of the rest of a noun phrase all at once, and produce a new restricted set from it. In this language, however, we don't have that luxury. If we want to keep things monocategorial, we need to find some way of keeping track of the base set and restrictions on the final quantified set simultaneously as additional quantifiers and other words are added in arbitrary orders. This will require altering our lexical semantics to account for a third argument, and that will in turn require altering the syntactic interpretation rules to provide that third argument.

The altered interpretation rules look like this:

[|S|] = ∃x.[|C|](x)

[|C: P C|] = λx. ∃y. ∃z. [|P|](x)(y)(z) & [|C|](x)
[|C: P .|] = λx. ∃y. ∃z. [|P|](x)(y)(z)

[|P: w P|] = λx.λy.λz.[|w|](x)(y)(z) & [|P|](x)(y)(z)

[|P: w ,|] = λx.λy.λz. [|w|](x)(y)(z) & y ⊆ z & ∃r. y ⊆ {z : ∃e. e ∈ x & r(e, z)}

Here we have simply asserted the existence of an additional set variable, z, and added a term to define y, the set of referents for a phrase, to be a subset of z.

The new forms of the different lexical semantic classes look like this:

a) λx.λy.λz. z  red
b) λx.λy.λz. y ⊆ {w : ∃e. e  x & ag(e, w)}
c) λx.λy.λz. x  run
d) λx.λy.λz. x = y
e) λx.λy.λz. |y| > |z - y|

Now, class-a noun-jective words operate on the set z, which forms the basis set for quantification. Relation words (classes b and d) act on y, which represents the actual referents of the phrase and is the result of quantification, and x, as before; and class-e quantifier words specify some relation between the set of referents y and the basis set z. The example given shows the semantics for the quantifier "most": the cardinality of the set of referents is greater than the cardinality of its difference with the basis set.

So far, we have not actually eliminated the need for logical existential quantifiers "built-in" to our semantics, but we have eliminated their semantic effect; all entities described by a sentence in the monocategorial language are no longer implicitly existentially quantified. Rather, the quantifiers in our predicate logic forms serve only to bind variables that we can use to refer to the different referent sets; and the referent sets are explicitly quantified by whatever quantifier words you feel like using. In the absence of any explicit quantifier word, sentences are evaluated as being true for some, unspecified, subset of the basis. This is exactly the same level of ambiguity present in natural languages (like Mandarin) which lack obligatory grammatical number.

Actually eliminating the existential quantifiers from our predicate logic forms would require directly constructing the relevant sets from unions and intersections, so as to eliminate the need for a common variable to use to tie the different parts together. That, in turn, requires separating the classes of quantifier and relation words from the class of noun-jectives, which cannot be done while maintaining the monocategorial analysis. It will, however, be possible to do so in WSL, which does have the necessary multiple syntactic levels.

It should also be noted that in adding quantifier words to the monocategorial language, we actually did not need to go all the way to introducing the full formalism of "generalized quantifiers"- and in fact, it looks like we can't, since the order of compositional operations which motivates the use of generalized quantifiers in the semantics of natural languages just doesn't exist here. The denotations of our monocategorial phrases are simple sets of referents, whose members have some thematic relation to the implicit event. In contrast, the denotations of natural-language noun phrases composed of nouns and generalized quantifiers are sets of possible sets of referents with the appropriately restricted cardinalities; and the correct set of referents is then extracted at the next higher level of composition, when a role is assigned to the noun phrase by a verb or adposition. Without that extra level, the monocategorial language must simply specify the final referent set directly. Again, however, WSL does have the more typical separation of nouns, quantifiers, and role-assigning words at different syntactic levels, and so we will be able to explore a more naturalistic analysis for that language.

For more thoughts on the monocategorial language, see this following post.

Monday, September 7, 2015

A Sister Language for WSL

My last post was triggered by recent discussions on CONLANG-L which inspired me to start formalizing the semantics of WSL. But after I started formalizing WSL, those discussions kept going. And as it happens, I got inspired to start on the design of a new language, similar in basic structure to WSL but also very different. So, before we get to Part II of WSL's semantic model, we're going to take a brief detour through the basic design of a new sister language.

Background: This idea came out of a discussion on how to describe the semantics of a monocategorial language whose complete syntax could supposedly be described by the following simple grammar:

 w S | 0

Or, "A sentence consists of a list of words." That's it. Any words in the language, in any order- all of them are grammatical sentences. Which, really, is equivalent to "no syntax at all". The obvious choice for semantic rules when presented with that syntax (or lack thereof) is David Gil's polyadic association operator, but the creator was adamant that that was not a correct analysis of his language. My conclusion was that " w S | 0" was simply not the correct grammar as claimed, but rather that it was a two-level structure that grouped words into distinct phrases, but where phrase boundaries are maximally ambiguous. This still permits every possible linear arrangement of words as a valid grammatical sentences, since the order of words within a phrase and the order of phrases within a sentence are still completely free.

With only one class of words and completely free word order, resulting in no way to tell where phrase boundaries are and how words should be grouped together, such a language would initially seem to be fairly useless- even discounting lexical ambiguity, the number of different possible interpretations grows as the square of the number of words in a sentence- an ambiguity load that dwarfs what exists in any natural language, and would quickly swamp what you can reasonably handle with pragmatics.

In a spoken language, though, more function words or morphological words would not necessarily be required to eliminate that ambiguity- phrase and sentence boundaries could be quite adequately delimited by intonation. And intonation can in turn by encoded in text via appropriate punctuation, while still reasonably claiming that this is a monocategorial language at the lexical level (although it will have multiple types of internal syntactic nodes). I've never really played with the intonation rules for a conlang before, and especially not the effect of intonation on semantics; and I haven't seen much of that documented in other people's conlangs, either. So, this is a pretty enticing opportunity to really isolate the semantics of suprasegmental intonation.

Now, the point of WSL was to create something that very obviously does not have anything that could reasonably be called a category of "verbs" at any level, but not necessarily to be simple or minimalistic. And WSL does in fact have quite an array of different parts of speech. But for this one, the aim will be to see how far it can go before it becomes necessary to add any additional lexical classes.

The syntax of this new language ends up looking up like this:

 C
 P C | P .
 w P | w ,

This reads as "A sentence consist of a clause, a clause consists of a phrase followed by another clause or a phrase followed by a period, and a phrase consists of a word followed by another phrase, or a word followed by a comma." We also specify the phonological / orthographical rule that a sequence of ", ." coalesces into a single "."

At the phonological level, the "," and "." are realized as particular intonation patterns on the preceding phrase. I'm not wedded to anything yet, but I'm thinking rising tone over the last word of a phrase for ",", and contrasting falling tone for ".". That would lead to an intonation pattern over a whole sentence that consists of a series of level tones followed by rises, and then terminated by a fall.

The extra level of rules that turns an S into a C may seem superfluous (and if we just want to describe syntactic structure by itself, they are), but the extra level makes the semantic interpretation rules much simpler.

Those basic interpretation rules look like this:

[|S|] = ∃x.[|C|](x)
"There exists some x such that the denotation of is true for x."

[|C: P C|] = λx. ∃y. [|P|](x)(y) & [|C|](x)
[|C: P .|] = λx. ∃y. [|P|](x)(y)
"For some x, there exists some y such that the denotation of P applied to x and y is true, and
the denotation of C is true for x."


[|P: w P|] = λx.λy.[|w|](x)(y) & [|P|](x)(y)
"For some x and y, the denotation of w and the denotation of P applied to x and y are true."

[|P: w ,|] = λx.λy. [|w|](x)(y) & ∃r. r(x,y)
"For some x and y, the denotation of w applied to x and y is true and some relation r exists between x and y."

Basically, this is just a fancy mathematically formalized way of saying that a sentence describes an event which gets passed into each sub-clause, and then each phrase describes its own separate entity, and the meaning of the whole sentence is just the conjunction of the meanings of each word, applied to the whole-sentence event and the entity for that word's containing phrase, along with the assertion that the entity for a phrase has some kind of relationship to the sentence.

Every word in the language has the "semantic interface" of a two-place predicate, or a two-argument curried lambda expression, taking in an event variable and an entity variable and specifying some restriction on either or both referents and/or a relationship between them.

Some words will be simple predicates that restrict the referent of the phrase, or tell you about its properties. They will have meanings the look something like this:

a) λx.λy. red(y)

which completely discards the event and just applies some predicate (in this case, "red") to the entity variable.

Some other words will be two-place relations that tell you about the thematic role of the entity in relation to the event. They will have meanings like

b) λx.λy. ag(x, y)

which tells you that the referent of this phrase (represented by the entity variable y) is the agent of the event.

And a third class of words will tell you about the event itself. These words could come in two sub-varieties; things that look like

c) λx.λy. run(x)

which discard the entity variable and just apply a predicate to the event; and things that look like

d) λx.λy. x = y

which tells you that the entity for this phrase is, in fact, an event, and that the event is a subset or superset of (or, in this particular case, simply is) the entity described by the enclosing phrase.

Now, semantic class c has the interesting property that, since the meanings of words in that class do not depend on the entity of the phrase, they can appear in any phrase in a sentence without altering the literal meaning. That's a fairly unique behavior, and could be used to argue for recognizing them as a separate part of speech from the rest, but they don't have to be analyzed so. Their syntactic behavior is undistinguished from every other word. Even so, I'm not sure if I will want to include some in the language for "fun", or if they should be disallowed so as to avoid the argument.

Also, the boundary between classes b and d is very fuzzy, since subset, superset, and identity could just as well be modeled as binary relations between a phrasal entity and a sentential event as things like "agent" and "patient" are.

Finally, the a category, which would typically seem to correspond with nouns and adjectives, also does not have any distinguished behavior compared to classes b, c, and d. Relation words and event words can be left out, and you can have a complete sentence that consists only of class-a semantic noun-jectives, which are asserted to exist and to have some unspecified relation to some unspecified event[1]. In WSL, role markers are obligatory, but here we have the extra "& ∃r. r(x,y)" in the interpretation of phrases just to account for the case where words with the semantics of a role marker are missing.
On the other hand, you can also leave out all class-a words, and have a complete sentence that consists only of class-b relations; and the same applies to the last two classes of event words as well. Finally, there are no selection rules that cause a word of any of the four classes to disallow the use of any other particular class in the same phrase or sentence; some combinations of words may be contradictory or nonsensical, but every string of words is grammatical, and can be interpreted.

It should also be possible to represent quantifiers in this framework, as totally undistinguished words at the syntactic level which merely happen to have another different internal structure in their lexical semantics. This would allow getting rid of some of the built-in existential quantifiers, but will first require removing a few layers of abstraction from my current semantic notation in order to uncover the set-theoretic mechanics of generalized quantifiers. My efforts to that effect are detailed in this follow-up post.

Next, I'd like to figure out some useful application for stress-marked focus, which could be indicated orthographically with Capital Letters or something. That will take some thinking, since English examples often rely on the semantics of some focus-sensitive lexical item, and using it that way would provide a good argument for recognizing focus-sensitive items as a second part-of-speech. But some really simple rising/falling intonation gets us pretty dang far doing nothing but marking linear phrase boundaries!

[1]  Which means that elliptical answers to questions aren't actually elliptical at all- they're still complete grammatical sentences!