Could humans learn a language made entirely of music?

Linguists explain how a language built from musical notes could convey complex meaning and even support conversation

An illustration of of two figures in profile facing one another with music notes and red line accents around them

oxygen/Getty Images;Scientific American Illustrations

Illustration of a Bohr atom model spinning around the words Science Quickly with various science and medicine related icons around the text

Rachel Feltman: For Scientific American’s Science Quickly, I’m Rachel Feltman. One quick heads up: This episode features some spoilers for Project Hail Mary.

If there’s one thing that the many, many viewers of the movie Project Hail Mary seem to agree on, it’s that the alien character known as Rocky is an absolute delight. In the adaptation of Andy Weir’s novel about a biologist-turned-teacher who finds himself burdened with a mission to save Earth, translation software gives Rocky a charmingly wry voice with which to communicate with the main character, Ryland Grace.

But while Rocky’s text-to-speech persona is used to great dramatic and comedic effect in the film, his native language is also worth a listen. This complex, music-based communication system wasn’t treated as an afterthought. Two experts in the field of constructed language actually, ya know, constructed his language. So what does it take to communicate using musical notes, and how would Rocky’s mother tongue actually work?


On supporting science journalism

If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


Here to tell us all about it is Allison Parshall, an associate editor for mind and brain at Scientific American.

Allison, thanks so much for joining us today.

Allison Parshall: Thanks for having me.

Feltman: So we are gonna talk about musical languages. We talk about music and language with you a lot, but today is different.

Parshall: [Laughs.] I’m always here for the same thing. But today is a different—it is quite literally musical languages. And I actually wanna start with a challenge for you. So I want you to pretend that you and I are in some sort of weird situation where we can only communicate using a piano. We can’t talk or write or use gestures or make any faces. We can just press the keys on the piano. Could we create a language that would allow us to carry a conversation like the one we’re having right now? Could we do that, and how would you do it?

Feltman: Well, I will say that I’m very excited for this Knives Out sequel that we’re going to be in together. Um, I guess my first move would be, like, literally using the letters, corresponding to the notes, in the scale as a code. You could probably get more letters in there if you made it more of a, like, alphanumerical thing. I think it could be done. I don’t think I could have a very good conversation that way, and I think it would make me mad to try.

Parshall: Right? Like I guess you could take the entire keyboard and give each letter of the English alphabet a note and then just play through. But that would be pretty time intensive if you think about it. Like just spelling out the word face, F-A-C-E. Like I can say “face” in milliseconds. It would take me quite a long time to go through as some sort of code.

So no human language operates like this based off of pitch alone. Like pitch is an important part of how we talk. We have tonal languages, like Mandarin and Yoruba, in which the pitch of a word or a syllable kind of distinguishes it from some others, and that pitch carries meaning.

And like, you know, in English I raise my pitch at the end of sentences to denote a question. But it’s not just based off of the pitch. We don’t have any languages where it’s just notes like what you’d play on a piano because, kind of, why would we? We have this beautiful apparatus of our vocal tract that has been shaped by evolution to kind of create meaning in very precise and quick ways. And even among people who make up languages for fun, they’re called constructed languages or conlangs, there aren’t many examples of languages that use pitch like this. But then enter Rocky the Eridian. Rachel, have you seen or read Project Hail Mary?

Feltman: I have not read it, but I did finally see it recently. I saved it to watch at home, and I was delighted.

Parshall: Can you give us a little recap of who Rocky is?

Feltman: Yeah. So Rocky is an alien who is a very alien alien, which is often a problem in sci-fi; the aliens are often pretty peoplelike. But Rocky is a rocky little dude who, like, chirps and sings to communicate, is great at building things, loves puppetry. Yeah, I have to say, I kept seeing headlines about how people were in love with Rocky, and I guess I’m so cynical now, I, like, don’t have much faith in the American consumer of film, that I was like, “Eh,” you know, “whatever.” And then I was like, “Oh, I’m also in love with Rocky. What a delightful little guy.” But yes, he ...

Parshall: You’re not above being in love with Rocky. No one’s above it.

Feltman: I was ready to be disappointed, but no, he is delightful, and yeah, he, like, uses sound as the primary sense and consequently has a very musical language.

Parshall: One thing that first endeared me—I read the book before I saw the movie—is that in the book, Rocky’s language is denoted with just little music note emojis. Like, they don’t, you know, they don’t specify the music notes. The audiobook has a little bit of music playing over it. But we learn that Rocky has five sets of vocal cords or vocal pipes that can each be played independently, and they all kind of come together with these pitches to create his language, which the human protagonist eventually learns. And I was reading this book both intrigued and admittedly skeptical that such a language could ever be understandable by humans. And in the movies, they were actually going to have to figure out how to do it. So I was delighted to learn that the creators of the movie hired prominent language creators Jessie Peterson and David Peterson. You’ve probably heard of at least one language they’ve created. Between them, they’ve created Dothraki and Valyrian from Game of Thrones. You can actually learn High Valyrian on Duolingo now. They’ve also worked on the Dune movies, on the new Supergirl and Superman movies. But obviously, like you said, the Eridian language is alien alien. It’s kind of another thing entirely. The final product is quite beautiful and complex. Here’s what it sounds like.

[CLIP of Rocky speaking in Project Hail Mary. Grace responds: “You want me to go back in my ship? But I just got here.”]

Feltman: You know what I like about it is that not only is it obviously a, a totally nonhuman language, but they didn’t even take the low-hanging fruit of like, “And the music sounds sad when he’s sad.” It’s just pretty. It’s just nice, nice sounding but does not sound like anything I could parse.

Parshall: Yeah, there’s almost like a whale song quality to it, kind of equally as alien. The sound design itself, like what the sounds coming out of his vocal pipes sound like, that actually came from another team, but the Petersons were in charge of, kind of, the theory behind it, sculpting, like, what an actual language would be that would work like this.

And I wanted to understand how the Petersons helped do it. So I spoke with them about their creative process, and here’s what Jessie Peterson, David Peterson’s, partner in language creation, had to say.

Jessie Peterson: It was very interesting because they, from the outset, had said they don’t want it to feel like a human language because he’s supposed to be really more musically communicating. He’s supposed to be using multiple airways or air paths, um, we refer to them as instruments, but that’s just sort of an easy way to think of it, sort of, internal instruments inside his body more than humans can do all at once.

Parshall: David’s interest in the theoretical possibility of musical language goes back decades.

David Peterson: This is something that I’ve been thinking about since I wanna say 2004 in, in grad school. I was talking about with a guy about creating a musical language with a guitar because it’s, like, the first thing that comes up is there’s a series of questions that you ask if you’re creating a musical language, or there are a series of constraints. So one is the ability for a human to understand it. So in other words, like: Do you have to have perfect pitch to be able to understand it, or is this something that the average person can understand? The next is: How much of the language can you convey? How robust is the lexicon going to be, or is it gonna be a reduced set? Then it’s: How easy is it for somebody to produce? In other words, do you have to be a serious musician to do this? The last one is: How nice is it to listen to? In other words, like, with sound, it’s like, you can do all kinds of things, but is it nice to listen to? And if it just sounds like garbage, then, kind of, what was the point of it?

Parshall: So the first musical language to be created, as far as the Petersons are aware, is called Solresol. It’s more than 200 years old. It was invented by this French music teacher, and it uses the notes of the major scale—you know, the [sings] do, re, mi, fa, sol, la, ti, do—as the syllables, kind of, of its language, which means it has very few syllables and kind of restricted meanings.

Let me see if I can do a couple of them: [Sings] “mi, sol” is good. And [sings] “do, mi, sol” is God, so you can see how they kind of relate to it. And then [sings] “sol, mi, do” is the devil. I guess they had their priorities straight in the, um, 1800s or 1700s in France.

It has a really reduced number of words that it can convey, or as linguists call it, a reduced lexicon. Another more recent musical conlang was more inspiring to the Petersons. It’s called Moss. It was created by a musician named Jackson Moore. It’s also very simple in some ways because it only has 120 words. I have a little sample for you of what a conversation sounds like.

[CLIP: “Conversations at La Mama,” by Jackson Moore]

David Peterson: It was very clever, because it actually sounded nice, and it was not too difficult, it was just notes. And of course, my original thought was like, if you’re just doing notes, you’re not gonna be able to get very many words. And so he leaned into that. He actually just reduced the lexicon radically and said, “Well, I’m only gonna do this many words.” Which is fine but not necessarily for, like, a spoken language that an actual being is gonna use. It has to be more robust than that.

I had come up with this idea at one point was: What if you had different people playing different instruments and this expanded the lexicon? With Rocky it was like, all right, this is one being that has all the instruments built in, and we can just assume that they’re good at parsing this stuff. So that gives us the full lexicon.

Parshall: So with Eridian biology, the Petersons had everything they needed to create a musical language that could be functional, beautiful, and convey an entire world of meaning with just a few musical notes. So in the book, Eridians have five sets of vocal cords, but the Petersons decide to pare it down to just three main vocal lines, each of them with a distinct-sounding tone, kind of like different instruments, that can each play one of seven specific notes.

So you got three lines, each of them have seven possibilities, each of them sound kind of different from each other, and they’re all playing over top of each other on the same beats. And that would give them just enough sound options to create a variety of meanings but still in a succinct way. So two of those three vocal lines are creating the meanings of the words themselves.

David Peterson: So for example, I don’t know, one, one of them—a lexical item might be, let’s say, two beats, and instrument one plays A-flat and C, and then instrument two plays A-flat and G, in that order and overlaid.

[CLIP: Music notes overlaid]

And there absolutely were things like affixation. So maybe a characteristic two-tone sequence would be a prefix or a suffix, and then it comes before or afterwards.

Parshall: The third line is where it gets really interesting. So as I talk, Rachel, you can tell where every one of my words starts and ends, right?

Feltman: Yes. [Laughs.]

Parshall: Even though I’m not, like, pausing between each word, you can still do that. And I’m curious if you have any theories about how you’re doing that.

Feltman: Yeah. So I know that this is something that I’ve, like, covered research on in the past. And if I’m remembering correctly, it’s like, there are structures in the brain that seem to be dedicated to this, and we’re, like, very tuned in to the context of the languages we know. And that, like, allows us to sort of parse very unconsciously what actually—there’s no, like, set rhyme or reason to, like, where these sounds go. There’s no, like, inherent wordiness to a word that, like, it makes sense that a brain should be able to parse, and yet we do, and that’s really cool.

Parshall: And yet we do.

Feltman: [Laughs.] And if you, if a language is totally foreign to you, then yeah, you might think you know where the words stop and end, but you may in fact be wrong.

Parshall: Yeah, actually, I, you know, listening to most languages that I don’t know, there’s no shot I could guess correctly where each word begins and ends. And you’re right, it’s kind of like our brain’s sophisticated pattern recognition picking up on patterns in rhythm, intonation, and importantly, kind of, the meaning embedded in the words themselves, how words tend to start and end and everything like that.

The example that David Peterson uses is: How do you know when I say the word “carpet” that I don’t mean a “car pet”? Like, you know, “car” and “pet.” Well, first of all, that’s just nonsensical. Second of all, there’s something—when I say “carpet,” I put a little bit of different sound on “-pit” instead of “-pet.” So you know, there’s all these little rules.

In Eridian, like English, the words vary in length; they could be one beat or multiple. So the Petersons decided to have this third vocal line dedicated just to conveying the, kind of, grammatical contours of the sentence. Like, the tone is going to be consistent when you’re in the same word, and then it’ll change when that word changes. So it’s kind of this convenient way built in to signal the different parts of a sentence. So what do you think, Rachel? If I were to put you in a spaceship 12 light-years away from Earth, and you only had Rocky for company, do you think you could learn and eventually fully understand his language?

Feltman: Absolutely not, but we’d be great friends anyway. We’d communicate via puppet, and that would be terrible ’cause I’m not good at sculpting, either. But no, in all seriousness, it does seem like a very tall order. And that’s as somebody who is musical; it’s hard to imagine, you know, where, where one would start.

Hearing how they use that third vocal line, I think I would get very into Eridian grammar. I was, like, really into Mandarin grammar when that was a language I was working on. I think it’s always really cool just seeing how different languages, you know, handle the need to, like, create syntax and structure in the way you put words together and the, sort of, words that are created just to, sort of, imbue meaning to a sentence. So listen, I would love to immerse myself in the process, but no, if we just had to understand each other, and there was nobody teaching me, I do not think it would happen.

Parshall: The weirdest immersion program of all time, yeah. I mean, there’s actually been some criticism of the movie and of the book that Grace and Rocky’s ability to understand each other kind of from zilch is pretty far-fetched. Some of that is just movie magic; you have to assume that they’d come to understand each other. But like, in the book, Grace stops using the translator pretty quickly. He just kind of is like, “Eh, I got it,” which I do think is pretty far-fetched. If only we all had that ability.

But for what it’s worth, I’m pretty similar to you. I think I could maybe figure it out eventually, but oh boy would there be a learning curve. David, hilariously—I asked him this question, and he was kind of like, “Yeah, I mean, sure. Why not? I could learn it.”

David Peterson: It’s not too tough. You get used to it, especially since it has a reduced set of notes, not the full set. That helps a little bit. I mean, I don’t think it would necessarily be the easiest thing in the world, but I think it would be possible.

Feltman: Well, I guess this language might be easier for some people to pick up than others. I mean, like I said before, I am a musical person, but I don’t have perfect pitch, by which I mean, I can’t, like, pull a note out of the air and have it match that note on a piano. So I imagine that would make this tougher for me?

Parshall: Yeah, like in our original little thought experiment of how you and I would communicate with a piano, we were kind of assuming that we would have to have perfect pitch, right? Like, if I want every single note on its own to carry meaning, I would have to be someone who could hear the note [hums a note] and know exactly what note that is. That’s a relatively uncommon ability among humans. What’s far more common is what I think you and I can probably both do, which is very easily trained. It’s called relative pitch. You can tell how far apart two notes are. Like, [hums two ascending notes] is that a fifth? I don’t know. Maybe it’s a fourth. [Laughs.]

That’s something that can be trained pretty easily, is telling how far apart two notes are. And Jessie says that while they assumed Eridians do have perfect pitch, they made the language so that you only need relative pitch in order to understand it.

Jessie Peterson: The way we wrote it, it was also—you don’t necessarily need to have perfect pitch. The important thing would be having the understanding of the relative difference between what was done before and what’s done next so that you could theoretically start on any note. But it’s like, what happens when you go from here to five steps down and three steps up, you know? So it’s more that rhythm of movement that would be necessary, and you would need to be able to hear and distinguish that, and hear and distinguish it in three different modes, plus, you know, the optional fourth and fifth instruments, which are adding things like surprise and disgust and, just, emotional tone to it. And so I think it would be difficult for a lot of human ears, but if you are exposed to it long term, I would imagine that even someone who struggles at the beginning would eventually get a feel for it to at least be able to probably get the gist of maybe what they’re hearing.

Parshall: The musical medium also let the Petersons suggest some really creative ways to convey tone and meaning based on how the words were delivered.

David Peterson: For example, if you think about a guitar, and you’re just playing a scale, the usual way is you just play the notes, you know, [sings] da, da, da, da, da, da, da, da. That was a little off. [Sings] da, da, da, da, da, da, da. I know that was off, too.

Parshall: It’s a mode. It’s okay.

David Peterson: Anyway, how one note flows into the other. But there’s also, like, you can do a staccato picking, where it’s like, [sings] bump, bump, bump, bump, bump, bump, bump, bump. And, you know, we said, like, maybe this could be done if, you know, Rocky is trying to be very precise, so, you know, [sings] ba, ba, bump, bump, bump, bump, bump, that type of thing. We also suggested something like hesitation could be expressed with, like, a tremolo thing, or you’d think, like, a whammy bar that like, [sings a scale in a warbled tone] that type of thing. And then the other one was like, since not every single note is meaningful, like, sarcasm could be expressed by going, like, you know, a half step off of each note. So if it’s like, [sings] da, da, dum, is one thing, it could be like, [sings.] da, da, da. You know, kind of off a little bit and be like, “Duh.”

Parshall: You even made that sound sarcastic.

David Peterson: So that type of a thing. We just kind of, like, threw the sound designers, like, these options saying, “These are things you can do.”

Parshall: So after talking to the Petersons, I kind of couldn’t believe that I ever doubted that creating a functional musical language was possible. Like, even though the Petersons spend most of their time creating languages that mimic how humans communicate, they and other conlang creators love pushing the boundaries of what’s possible.

David Peterson: Once you kind of set your sights on, well, “What might be a possible way to convey a message?” not, “What is a usual way?” or “What would be a convenient way necessarily for two humans to communicate?” If you just say, “What’s a possible way?” There are so many possibilities that haven’t been breached. Like we have a conference we run, and the first one we did, Kopikon, I created a language whose medium was Jelly Belly jelly beans.

Parshall: What?

David Peterson: So like, you would just take a handful of jelly beans, and the flavors would give you a message. And they didn’t need to be taken in order ’cause that obviously wouldn’t be any fun.

I discovered—in fact, all of us discovered who were there in the audience—that there is a boundary that you hit when it comes to human tongues. Humans aren’t great at making these distinctions of taste, especially when there’s something that’s really overpowering that dulls your taste buds. But it’s like, this is something that you could do. There are lots of possibilities that haven’t been broached yet in terms of creating a language.

Parshall: So yeah, don’t let your dreams be dreams, kids. Make a musical language, Jelly Belly language—your tongue’s the limit.

Feltman: Yeah, that’s incredible. And, you know, I think a lot of people were so taken with Rocky as a character and really enjoyed the scenes after he gets his, you know, computerized voice, which is very cute. You know, all the “amaze, amaze, amaze,” it was very well done. But I hope that, you know, people will rewatch this with an increased appreciation for the beauty of his own mother tongue.

Parshall: Hmm. And I would love the opportunity to have a very strange language immersion program with Rocky the Eridian.

Feltman: Absolutely. Well, Allison, thanks so much for coming on to share this with us. It was a delight.

Parshall: Thanks for having me. And the next time there is music and language news, you will be the first to know.

Feltman: [Laughs.]

That’s all for today’s episode. We’ll be back on Friday for, coincidentally, another fascinating story about musical communication. Join us as we learn how nonverbal actors are using AI-powered tools to perform in a new opera production.

Science Quickly is produced by me, Rachel Feltman, along with Fonda Mwangi and Jeff DelViscio. This episode was reported by Allison Parshall and edited by Alex Sugiura. Marielle Issa and Aaron Shattuck fact-check our show. Our theme music was composed by Dominic Smith. Subscribe to Scientific American for more up-to-date and in-depth science news.

For Scientific American, this is Rachel Feltman. See you next time!

Rachel Feltman is former executive editor of Popular Science and forever host of the podcast The Weirdest Thing I Learned This Week. She previously founded the blog Speaking of Science for the Washington Post.

More by Rachel Feltman

Allison Parshall is associate editor for mind and brain at Scientific American and she writes the weekly online Science Quizzes. As a multimedia journalist, she contributes to Scientific American's podcast Science Quickly. Parshall's work has also appeared in Quanta Magazine and Inverse. She graduated from New York University's Arthur L. Carter Journalism Institute with a master's degree in science, health and environmental reporting. She has a bachelor's degree in psychology from Georgetown University.

More by Allison Parshall

Fonda Mwangi is an award-winning senior multimedia editor at Scientific American and showrunner of Science Quickly. She previously worked at Axios, the Recount and WTOP News. She holds a master’s degree in journalism and public affairs from American University in Washington, D.C.

More by Fonda Mwangi

Alex Sugiura is a Peabody and Pulitzer Prize–winning composer, editor and podcast producer based in Brooklyn, N.Y. He has worked on projects for Bloomberg, Axios, Crooked Media and Spotify, among others.

More by Alex Sugiura

It’s Time to Stand Up for Science

If you enjoyed this article, I’d like to ask for your support. Scientific American has served as an advocate for science and industry for 180 years, and right now may be the most critical moment in that two-century history.

SciAm always educates and delights me, and inspires a sense of awe for our vast, beautiful universe. I hope it does that for you, too.

If you subscribe to Scientific American, you help ensure that our coverage is centered on meaningful research and discovery; that we have the resources to report on the decisions that threaten labs across the U.S.; and that we support both budding and working scientists at a time when the value of science itself too often goes unrecognized.

In return, you get essential news, captivating podcasts, brilliant infographics, can't-miss newsletters, must-watch videos, challenging games, and the science world's best writing and reporting. You can even gift someone a subscription.

There has never been a more important time for us to stand up and show why science matters. I hope you’ll support us in that mission.

Thank you,

Jeanna Bryner, Editor in Chief, Scientific American

Subscribe