MSc Applied Cognitive Psychology — Utrecht University (Incoming) BSc Psychology BA Business Administration

Berke Tan Tabak

I design behavioral studies about how people judge technology, especially where trust, creativity, and music meet. Then I turn the findings into questions that product teams can act on.

See my work below ↓
Featured Case Study

How AI Labels Change the Way People Evaluate Music

For my undergraduate thesis, I built a controlled behavioral experiment with an fNIRS layer to study how AI attribution changes the way people experience music.

As an electronic music producer and artist, I became interested in the rapid integration of AI-generated music into everyday listening environments.

Generative systems were appearing across streaming platforms, festivals, social media, and clubs, while listeners often had limited information about how the music had been created.

I wanted to understand what happens when platforms distinguish between "Human Composed", and "AI-Generated" creative work — and whether attribution changes the value listeners assign to the same listening experience.

What happens to the listener’s perception when the origin of the music becomes visible?
Undergraduate thesis · independently developed and executed Presented June 3, 2026 Supervisors: Prof. Dr. Seda Dural and Özce Cinoğlu Jury evaluation · 98/100
My roleResearcher and sole project owner

I developed and executed the project myself, from the original question and study design to programming, data collection, analysis, and presentation.

Study design2 × 2 × 3 within-subject experiment

Displayed label × actual producer × musical genre.

MethodsBehavioral measures + fNIRS

PsychoPy task, five listener ratings, repeated-measures analysis, and prefrontal HbO.

What it becameA research case for product teams

I translated the findings into practical questions about labeling, trust, and platform testing.

01

Why I studied this

Much of the conversation around AI music seemed to be driven by either excitement or fear. I wanted to put one part of that debate under controlled conditions: could a source label change how listeners valued the same musical experience?

Three things mattered to me:

Take artists' concerns seriously by measuring how attribution affects the perceived value of creative effort.
Give listeners a clearer basis for informed judgment about authorship and different forms of human–AI collaboration.
Give streaming services, labels, and technology providers evidence they could use when designing labels, creator tools, and trust strategies.
02

What I tested

I designed a within-subject experiment that separated what a track actually was from what the listener was told it was.

Participants listened to 12 musical excerpts. Within each genre, tracks included both human-composed and AI-generated music, while the displayed source label was independently manipulated as either “AI-generated” or “human-composed.”

ArabesqueA culturally embedded condition for the Turkish sample, associated with lived experience, identity, and emotional expression.
BluesA culturally external, acoustic condition associated with musicianship, craft, and human performance.
ElectronicA technology-native condition in which digital production is already expected and may feel more compatible with AI attribution.

This design allowed me to test whether evaluation followed the sound itself, the displayed label, or the cultural and production context in which that label appeared.

What I measured — and why
Emotional InvestmentHow much emotion listeners believed the creator had put into the piece.
Originality / AuthenticityHow original the piece seemed; reported as “Authenticity” in the thesis materials.
QualityHow high in quality the piece seemed, beyond whether it was simply enjoyable.
LikingImmediate aesthetic enjoyment — used to distinguish pleasure from higher-order judgment.
TrustHow strongly listeners trusted that the piece had been produced by the displayed source.
Prefrontal HbO via fNIRSAn exploratory neural layer recorded after the source label appeared.
03

Why it matters

AI music labeling is not only a technical or regulatory question.

It affects how listeners understand authorship, how creators perceive the value of their effort, and how platforms mediate trust between human artists, generative systems, and audiences.

Carefully designed labels could support informed choice while giving streaming services, labels, and technology providers a stronger evidence base for product strategy, creator tools, and audience trust.

Poorly designed labels could mislead users, distort perceived creator value, or collapse very different forms of AI involvement into a single category.

How the Experiment Worked

Music first. Source label second.

Thirty participants first heard each track without a source label. After 10 seconds, the displayed label appeared and the track continued before ratings were collected. Across the study, displayed label, actual producer, and genre were varied in a fully randomized and counterbalanced 2 × 2 × 3 repeated-measures design.

Step 01

Listen without a label

Each trial began with an unknown track. Participants listened without being told whether the music was created by a human or by AI.

Step 02

Reveal the source

After 10 seconds, the track was labeled as either “AI-generated” or “human-composed.”

Step 03

Continue listening

The music continued after the label appeared, allowing attribution to shape the remaining listening experience.

Step 04

Rate the experience

Participants rated liking, originality/authenticity, quality, emotional investment, and trust in the displayed source.

Step 05

Compare sound and attribution

Because displayed label and actual producer varied independently, the experiment could separate reactions to the music from reactions to its perceived source.

Evidence hierarchy: Behavioral outcomes formed the primary evidence base. Prefrontal activity was recorded with fNIRS as an exploratory layer and is interpreted cautiously because only 13 participants remained after channel-quality exclusion.

Isolate attribution

Revealing the label after 10 seconds let participants encounter the sound before receiving information about its source. Ratings were collected after the labeled listening period.

Separate label from reality

Crossing displayed label with actual producer allowed the study to test framing effects without assuming that AI- and human-produced audio would differ acoustically in a consistent way.

Match claims to evidence

Strong behavioral patterns are presented as primary findings; the reduced fNIRS sample is treated as a direction for replication rather than proof of a neural mechanism.

Experiment buildPsychoPy 2023.2.3 · Pygame · Psychtoolbox · Lab Streaming Layer
Behavioral analysisIBM SPSS Statistics 29 · Repeated-measures GLM · Bonferroni comparisons
NeurophysiologyfNIR Devices 1000 · COBI Studio · Custom Python preprocessing and ROI analysis
Key Findings

There was no single, universal “AI bias.”

The label changed higher-order judgments more than basic enjoyment, and its meaning shifted across genres. The patterns below describe this sample and provide hypotheses for larger, cross-cultural platform studies.

Mean ratings by displayed source label

Each colored shape represents one genre. A point farther from the center means a higher average rating. Compare the same color across the two panels to see what changed when the displayed label changed.

N = 30 Repeated-Measures GLM 7-point Likert
Arabesque Blues Electronic
AI Generated0246Emotional Inv.TrustAuthenticityLikingQuality
Human Composed0246Emotional Inv.TrustAuthenticityLikingQuality
Exact displayed-label × genre means from the cleaned N=30 behavioral dataset, redrawn as responsive vector graphics in the original visual system.
Perceived Value

People did not stop liking the music — but they valued it differently.

Across conditions, Human labels increased perceived creator emotion by 0.58 points, perceived originality by 0.38, and quality by 0.32 on seven-point scales. Liking was identical at 4.44 under both labels. The attribution changed judgments about creative value without changing immediate enjoyment.

Creator emotion p=.002, η²p=.276 · Perceived originality p=.002, η²p=.285 · Quality p=.007, η²p=.227 · Liking p=1.000
Cultural & Production Context

The AI label did not mean the same thing in every musical context.

The three genres were deliberately selected to represent contrasting relationships with cultural familiarity, acoustic human performance, and technology-led production.

Arabesque · Culturally embedded Human-attribution premium

Human labels increased attributed creator emotion, perceived originality, quality, and confidence that the displayed source produced the piece.

Blues · Culturally external & acoustic Comparatively stable judgments

Ratings across the five measures changed comparatively little between the two displayed source labels.

Electronic · Technology-native A split response

Human labels increased perceived creative value, while the AI label was considered more credible as the stated source.

These patterns suggest that responses to AI attribution depend on both the label and the cultural and production context in which it appears.
Label × Genre

Trust reversed with musical context.

The label had no overall effect on source-claim trust (AI M=4.62; Human M=4.59). Averaged across actual producer, a Human label increased trust in Arabesque by 1.13 points, made no reliable difference in Blues, and reduced trust in Electronic by 1.32 points. Here, trust means confidence that the displayed source really produced the piece—not general trust in AI.

Human − AI label, Holm-adjusted: Arabesque +1.13, p=.003 · Blues +0.10, p=.632 · Electronic −1.32, p=.004 · Label × Genre p<.001, η²p=.339
Meaning vs Enjoyment

Listeners attributed the most creator emotion to Arabesque — but liked Blues most.

Arabesque received the highest ratings for emotion attributed to the creator (M=4.86) and perceived originality (M=4.15), while Blues received the highest liking (M=5.00) and quality ratings (M=5.06). In this sample, perceived creative meaning did not map directly onto enjoyment or perceived craft.

Genre: Creator emotion p<.001, η²p=.587 · Perceived originality p<.001, η²p=.290 · Quality p<.001, η²p=.297 · Liking p=.011, η²p=.165 (Greenhouse–Geisser corrected)
Actual Source

Displayed attribution mattered even when actual producer was controlled.

The experiment varied the displayed label independently from who actually produced the track. Actual producer showed no significant main effect across the five ratings, while displayed attribution shaped perceived value and interacted with musical context.

Actual producer: Creator emotion p=.162 · Perceived originality p=.491 · Source-claim trust p=.620 · Liking p=.261 · Quality p=.292
Exploratory Measurement

The neural layer generated a follow-up hypothesis — not a product claim.

The retained fNIRS sample suggested a possible right-prefrontal difference after AI labels. Because the evidence was not sufficient to establish a neural mechanism, product implications on this page are grounded in the behavioral findings.

Exploratory signal · interpreted separately from the primary behavioral evidence
Complete statistical record Behavioral repeated-measures GLM results

The cards above prioritize interpretation. This record keeps the behavioral tests visible for readers who want to inspect the evidence. Effect sizes are partial eta squared (η²p). When Mauchly’s test indicated a sphericity violation, the Greenhouse–Geisser-corrected p value is reported.

Displayed label effects

Emotional InvestmentAI label lower · p=.002 · η²p=.276
Originality / AuthenticityAI label lower · p=.002 · η²p=.285
QualityAI label lower · p=.007 · η²p=.227
LikingNot significant · p=1.000 · η²p=.000
Trust in stated sourceNot significant · p=.872 · η²p=.001

Displayed Label × Genre

Emotional Investmentp=.039 · η²p=.106
Originality / Authenticityp=.031 · η²p=.113
Trust in stated sourcep<.001 · η²p=.339
Qualityp=.014 · η²p=.137
LikingNot significant · p=.215 · η²p=.052 · Greenhouse–Geisser corrected

Genre effects

Emotional Investmentp<.001 · η²p=.587
Originality / Authenticityp<.001 · η²p=.290
Qualityp<.001 · η²p=.297
Likingp=.011 · η²p=.165 · Greenhouse–Geisser corrected
Trust in stated sourceNot significant · p=.663 · η²p=.014

Actual Producer × Displayed Label

Emotional Investmentp=.049 · η²p=.127
Originality / AuthenticityNot significant · p=.885 · η²p=.001
Trust in stated sourceNot significant · p=.050 · η²p=.126
LikingNot significant · p=.249 · η²p=.046
QualityNot significant · p=.357 · η²p=.029

Actual Producer

Emotional InvestmentNot significant · p=.162 · η²p=.066
Originality / AuthenticityNot significant · p=.491 · η²p=.016
Trust in stated sourceNot significant · p=.620 · η²p=.009
LikingNot significant · p=.261 · η²p=.043
QualityNot significant · p=.292 · η²p=.038

Additional interaction results

Emotion: Producer × Genrep=.040 · η²p=.105
Emotion: Producer × Label × GenreNot significant · p=.054 · η²p=.096
Originality: Producer × GenreNot significant · p=.101 · η²p=.076
Originality: Producer × Label × GenreNot significant · p=.195 · η²p=.055
Trust: Producer × GenreNot significant · p=.479 · η²p=.025
Trust: Producer × Label × Genrep=.002 · η²p=.187
Liking: Producer × GenreNot significant · p=.514 · η²p=.023
Liking: Producer × Label × GenreNot significant · p=.541 · η²p=.021
Quality: Producer × Genrep=.020 · η²p=.140 · Greenhouse–Geisser corrected
Quality: Producer × Label × Genrep=.018 · η²p=.129

Follow-up label effects within genre

OutcomeArabesqueBluesElectronic
Emotional Investment+0.88 · p=.014+0.07 · p=.800+0.78 · p=.003
Originality / Authenticity+0.53 · p=.026−0.07 · p=.747+0.68 · p=.006
Trust in stated source+1.13 · p=.003+0.10 · p=.632−1.32 · p=.004
Liking+0.32 · p=.418−0.07 · p=.724−0.25 · p=.629
Quality+0.45 · p=.035−0.08 · p=.587+0.58 · p=.012
Arabesque trustHuman − AI: +1.13
Blues trustHuman − AI: +0.10
Electronic trustHuman − AI: −1.32
Arabesque emotionAI-label penalty: −0.88

Behavioral N=30. Positive follow-up contrasts indicate higher ratings under the Human label; negative contrasts indicate higher ratings under the AI label. Follow-up p values come from paired comparisons averaged across actual producer and are Holm-adjusted across the three genres within each outcome. “Trust” refers to confidence that the piece was produced by the displayed source. fNIRS is intentionally excluded from this behavioral statistical record and remains an exploratory layer.

From Evidence to Product Decisions

Treat AI labeling as a product system—not only compliance copy.

The study does not prescribe one universal label. It identifies the decisions a platform should validate before deploying AI attribution at scale.

Priority 01

Define levels of AI involvement

Separate fully generated, AI-assisted, and human-led creation. A binary AI/human label risks hiding the contribution users actually want to understand.

Priority 02

Validate wording and genre together

Blues was comparatively label-resistant, while Arabesque and Electronic were sensitive in different ways. Genre should be treated as an experience variable—not background metadata.

Priority 03

Measure creative value and behavior together

Track originality, attributed emotion, quality, and source-claim trust alongside skips, replays, saves, shares, and creator support. Liking alone missed the most important effects.

Recommended next study

Test label taxonomy in a real streaming interface.

Compare “AI-generated,” “AI-assisted,” “human–AI collaboration,” and no-label conditions with a larger, cross-cultural sample. Randomize label wording and genre, add a manipulation check, and evaluate both perception and downstream behavior.

TrustAuthenticitySkip rateReplaySaveShareArtist followCreator support
Scope & Next Steps

Strong behavioral signals, with clear boundaries.

The study identifies where labeling changed perception in this sample. It is evidence for what platforms should test next—not a universal rule for every listener, genre, or interface.

Audience

Broader listeners are the next test

The behavioral sample comprised 30 Turkish university students. A larger, cross-cultural study should test whether the genre patterns travel across listener groups and markets.

Neural layer

Exploratory, not the basis of the recommendations

The fNIRS layer retained 13 participants after quality screening. It is treated as a preliminary signal; the product implications on this page are grounded in the stronger behavioral evidence.

Application

Move from perception to behavior

The controlled task measured judgments. The next study should place labels inside a real streaming interface and test skips, replays, saves, shares, follows, and creator support.

About Me

I move between the lab, the product, and the studio.

I have never experienced psychology, business, and music as separate paths. Psychology helps me ask better questions about people. Business keeps me focused on the decisions that research should inform. Producing and performing music gives me first-hand knowledge of creators, audiences, and the emotional stakes behind creative technology.

How I like to work

I am most engaged when a problem is still unclear and needs to be turned into something testable. I enjoy building the study, working through the messy parts of the data, and then explaining the result without hiding its limitations.

This thesis began with a question I genuinely cared about as a producer. It became an experiment because I wanted evidence, not just an opinion. That is also the kind of work I want to continue doing: research grounded in real human concerns and useful to the people making products.

EducationBSc Psychology, High Honors · BA Business Administration, Honors — Izmir University of Economics · Incoming MSc Applied Cognitive Psychology at Utrecht University
Product contextCoordinated product activities and stakeholder communications for Kuario, a Dutch payment and printing platform, at Oxivo International.
MusicElectronic music producer, DJ, and instructor under CHOSSN