Mathematics · 14 May 2026

Whose LLM Is It Anyway? Linguistic Comparison and LLM Attribution for GPT-3.5, GPT-4 and Bard

Ariel Rosenfeld, Teddy Lazebnik

Affiliations
  1. Department of Information Science, Bar Ilan University, Ramat Gan 5290002, Israel; ariel.rosenfeld@biu.ac.il
  2. Department of Information Systems, University of Haifa, Haifa 3498838, Israel
  3. Department of Computing, Jonkoping University, SE-551 11 Jonkoping, Sweden
ACML authorsTeddy LazebnikPI

The paper at a glance

Large language models can write text that matches or surpasses human quality, but do they have recognizable styles, as human authors do? We compared the vocabulary, parts of speech, sentence structure and sentiment of texts that GPT-3.5, GPT-4 and Bard produced for diverse inputs. We found significant linguistic differences, which allowed a simple off-the-shelf classifier to identify which model wrote a given text with 88% accuracy.

88%accuracy attributing a text to its source LLM

Key findings

  • Texts from GPT-3.5, GPT-4 and Bard showed significant differences in vocabulary, part-of-speech distribution, dependency distribution and sentiment.
  • These differences let a simple off-the-shelf classification model attribute a text to its source model with 88% accuracy.
Figure 1. Part-of-Speech (POS) distribution comparison between GPT-3.5, GPT-4, and Bard.
Figure 1. Part-of-Speech (POS) distribution comparison between GPT-3.5, GPT-4, and Bard. See it in the paper
On this page
  1. Abstract
  2. 1. Introduction
  3. 2. Methods and Materials
  4. 2.1. Data
  5. 2.2. Analytical Approach
  6. 3. Results
  7. 3.1. Vocabulary
  8. 3.2. Part-of-Speech (POS)
  9. 3.3. Dependency Parsing
  10. 3.4. Sentiment Analysis
  11. 3.5. Baseline and Ablation Analysis
  12. 3.6. LLM Attribution
  13. 4. Discussion
  14. Appendix A
  15. Appendix A.2. Additional Results
  16. Article notes
  17. References

Abstract

Large Language Models (LLMs) are capable of generating text that is similar to or surpasses human quality. However, it is unclear whether LLMs tend to exhibit distinctive linguistic styles akin to how human authors do. Through a comprehensive linguistic analysis, we compare the vocabulary, Part-of-Speech (POS) distribution, dependency distribution, and sentiment of texts generated by three of the most popular LLMS today (GPT-3.5, GPT-4, and Bard) to diverse inputs. The results point to significant linguistic variations which, in turn, enable us to attribute a given text to its LLM origin with a favorable 88% accuracy using a simple off-the-shelf classification model. Theoretical and practical implications of this intriguing finding are discussed.

1. Introduction

Large Language Models (LLMs), such as GPT-3.5 [1], GPT-4 [2] and Bard [3], have revolutionized and popularized natural language processing and AI, demonstrating human-like and super-human performance in a wide range of text-based tasks [4–7]. While the layman may find the responses of LLMs hard to distinguish from human-generated ones [8,9], a plethora of recent literature has shown that it is possible to successfully discern human-generated text from LLM-generated text using various computational techniques [10–12]. Among the developed techniques, the linguistic approach, which focuses on the structure, patterns, and nuances inherent in human language, stands out as a promising option that offers both high statistical performance [13] as well as theoretically-grounded explanatory power [14], as opposed to alternative “black-box” machine learning techniques [15,16].

Indeed, recent literature has shown that human and LLM-generated texts are, generally speaking, linguistically different across a wide variety of tasks and datasets including news reporting [14], hotel reviewing [17], essay writing [13] and scientific communication [18] to name a few. Common to these and similar studies is the observation that LLM-generated texts tend to be extensive and comprehensive, highly organized, follow a logical structure or are formally stated, and present higher objectivity and lower prevalence of bias and harmful content compared to human-generated texts [19].

Extensive research into human-generated texts has consistently demonstrated the inherent diversity in human writing styles, resulting in distinct linguistic patterns, structures, and nuances [20–22]. Notably, highly successful techniques for author attribution [23,24] and author profiling [25,26] have leveraged linguistic markers to identify and differentiate between authors and their characteristics (sometimes referred to as Stylometrics [27]). These remarkable capabilities underscore both the richness and variability present in human-generated texts as well as the unique linguistic traits presented, rather consistently, by different authors. Unfortunately, to the best of our knowledge, a similar inquiry into LLM-generated texts has yet to take place. That is, it remains unclear whether different LLMs present distinct linguistic styles and, if so, whether these linguistic markers could be effectively used for LLM attribution (i.e., identifying which LLM has generated a given text).

In this work, we report on a comprehensive linguistic comparison of LLM-generated texts generated by three of the most popular LLMs today: GPT-3.5, GPT-4, and Bard. Using a wide range of topics and prompts, the results reveal that, indeed, these LLMs are linguistically different, particularly in terms of vocabulary, Part-of-Speech (POS), and dependencies. In turn, these linguistic markers are shown to bring about a remarkable performance when applied to the LLM attribution task.

2. Methods and Materials

2.1. Data

We build upon the highly influential Human ChatGPT Comparison Corpus (HC3) [28]. HC3 is the most utilized database today for comparing LLM to human-generated texts. In its English portion, it encompasses both human and GPT-3.5 responses (termed ChatGPT in the original manuscript) to identical inputs divided into five distinct datasets: Finance [29], Medicine [30], Long-Form Question Answering [31] (aka reddit_eli5), Open-Domain Question Answering (aka open_qa) [32], and Computer Science (aka wiki_csai dataset) [28]. Here, we consider 1000 randomly sampled inputs from each dataset, along with their GPT- 3.5 responses. We then extend the database to include the matched responses provided by GPT-4 and Bard to these inputs. The GPT-3.5 responses were taken directly from the original HC3 corpus [28]. The additional GPT-4 and Bard responses were generated by submitting the same sampled HC3 inputs to each model. GPT-4 was accessed via the OpenAI API. Bard was accessed via the Google Bard web interface with code-based automation. No model-specific prompt adaptation was performed. The resulting database, which we term the LLM Comparison Corpus (LC2), consists of 5000 inputs and 15,000 responses (5000 from each examined LLM). LC2, along with our entire code. For reproducibility, the same prompting protocol was used across the examined LLMs. Specifically, each model was provided with the corresponding HC3 input using the prompt template reported in Appendix A.1. No model-specific prompt adaptation was performed.

2.2. Analytical Approach

Drawing upon the comprehensive linguistic analysis of HC3 [28], we analyze the key linguistic features of LC2 as follows: we examine possible linguistic differences in terms of vocabulary, POS, dependencies, and sentiment across the three LLMs (see [33] for a linguistic overview of these and related concepts). For clarity, we briefly define the main NLP concepts used throughout the analysis. Part-of-Speech (POS) tagging refers to the process of assigning each token in a text a grammatical category, such as noun, verb, adjective, adverb, or pronoun [34]. Dependency parsing, in contrast, analyzes the grammatical relations between tokens in a sentence [35]. In this framework, a sentence is represented as a directed graph: tokens (i.e., parts of words or full words) are represented as nodes, and directed edges indicate syntactic relations between them, such as nominal subject, direct object, adjectival modifier, or auxiliary verb. Thus, POS tagging characterizes individual tokens, whereas dependency parsing characterizes the syntactic relations among them. Following common practice in modern dependency-parsing systems, punctuation marks are also treated as tokens and may therefore appear as nodes in the dependency graph, with corresponding punctuation dependency relations. Finally, sentiment analysis refers to the automatic classification of the affective orientation of a text, commonly into categories such as positive, neutral, or negative sentiment [36].

Statistically, for comparing vocabulary, we use an ANOVA test with Tukey post hoc pairwise t-testing. Similarly, for comparing the distributions over POS and dependencies, we use Kolmagorov–Smirnov testing with the Bonferroni correction. For sentiment analysis, given its ordinal nature, we use a Wilcoxon signed-rank test. Notably, in our setting, sentiment analysis is used as a coarse linguistic marker capturing whether a generated response is predominantly positive, neutral, or negative in tone. Although sentiment is less syntactic than POS and dependency features, it captures another potentially relevant stylistic dimension of LLM-generated text, namely the affective orientation expressed in the response. Statistical significance is set to 0.05. Finally, we use the aforementioned linguistic markers as input to an off-the-shelf supervised machine learning model (XGBoost [37]) for the task LLM attribution—i.e., classifying a given text to its assumed LLM origin. The model’s performance is reported in standard form.

3. Results

3.1. Vocabulary

We start by examining the vocabulary presented by the studied LLMs. We focus on the following three characteristics: Average length (L)—the average number of words in each response; Vocabulary size (V)—the total number of unique words used in all responses; and Density (D)—which is defined as

D = 100V L · N

where N is the number of responses in the relevant dataset. The results are summarized in Table 1.

Table 1. Average response length, vocabulary size, and density of GPT-3.5, GPT-4, and Bard using five datasets.
DatasetLLMAverage LengthVocabulary SizeDensity
GPT-3.5208.1320,9742.49
FinanceGPT-4197.5322,7852.73
Bard219.2821,8092.64
GPT-3.5206.1479103.11
MedicineGPT-4168.0988275.69
Bard180.1675943.24
GPT-3.5142.6115,3799.06
open_qaGPT-488.4212,09716.93
Bard65.7410,82917.34
GPT-3.5191.3845,1981.40
reddit_eli5GPT-4151.1848,0952.05
Bard133.7046,1471.87
GPT-3.5202.3993475.03
wiki_csaiGPT-4215.0510,0746.73
Bard186.1892407.18

Starting with the average length, the three LLMs are found to statistically differ from one another. Specifically, Bard provides the shortest average responses. In fact, it provides the shortest average answers in three out of the five datasets. In the remaining two datasets, Bard provides the longest average responses in one case and the second longest in the other. In addition, GPT-3.5 provides longer average responses than GPT-4 (i.e., in four of the five datasets). Turning to vocabulary size, once more, Bard provides the smallest number of unique words in three out of the five datasets whereas in the remaining two, it comes second to GPT-4. Interestingly, GPT-4 provides greater vocabulary size than GPT-3.5 in four of the five settings. When considering density, GPT-4 demonstrates the highest density metric in three of the five datasets whereas in the remaining two datasets, Bard brings about a higher density.

Taken jointly, Bard seems to provide shorter responses compared to GPT-3.5 and GPT- 4 while also presenting smaller vocabulary size and relatively high density. In addition, GPT-4 seems to provide shorter responses than GPT-3.5 while presenting higher vocabulary size and density.

3.2. Part-of-Speech (POS)

Next, we examine the POS distribution of the provided responses (the complete list, along with their abbreviations and complete results, are provided in Appendix A). The results are summarized in Figure 1. The three LLMs provide statistically different POS distributions.

Part-of-Speech (POS) distribution comparison between GPT-3.5, GPT-4, and Bard
Figure 1. Part-of-Speech (POS) distribution comparison between GPT-3.5, GPT-4, and Bard.

We start by examining the use of nouns and adjectives, the most prevalent POSs for all three LLMs. Starting with nouns, we see that GPT-3.5 uses nouns more often than the other LLMs both in singular (NN: 30.36% vs. 28.73% by GPT-4 and 26.15% by Bard) and plural form (NNS: 16.19% vs. 13.51% by GPT-4 and 12.67% by Bard). In turn, GPT-4 uses nouns more often than Bard in both forms, although the differences are much smaller. That is not the case for singular-form proper nouns (NNP), which are around twice less prevalent than non-proper nouns (NN), where GPT-4 demonstrates higher use patterns than Bard (11.43% vs. 9.91%) which, in turn, uses them more often than GPT-3.5 (8.55%). When plural-form proper nouns (NNPS) are considered, all LLMs scarcely use them (all are below 0.3% prevalence). However, GPT-3.5 uses them almost two times less often than the others (0.17% compared to 0.29% and 0.26% by GPT-4 and Bard, respectively). Turning to adjectives (JJ) and the significantly less prevalent comparative adjectives (JJR), once more, we see that GPT-3.5 provides the greatest prevalence followed by GPT-4 and Bard (JJ: 15.12%, 14.15%, 12.75%; JJR: 0.72%, 0.68%, 0.89%). When superlative adjectives (JJS) are concerned, all LLMs rarely use them (all are below 0.9%). However, Bard uses them more often as the GPT-4 which, in turn, uses them more than GPT-3.5 (0.82% vs. 0.26%, 0.19%).

Considering verbs and adverbs, the results point to very minor differences between the LLMs. Specifically, GPT-3.5 slightly uses present tense verbs (VBG and VBP) more often than the others whereas Bard uses past verbs (VBN and VBD) and third-person verbs (VBZ) slightly more often. In turn, GPT-4 presents a marginally higher use of base-form verbs (VB) than the others (for the complete results, see Appendix A). Turning to adverbs, GPT-3.5 presents a lower use of adverbs (3.66%) compared to GPT-4 and Bard (4.87% and 4.15%, respectively).

The remaining POSs (7.26–23% of the text), encompass a relatively small portion of the text. However, considerable differences between the LLMs are encountered. Specifically, Bard presents a relatively high use of personal pronouns (PRP: 1.7% vs. 0.69% by GPT-3.5 and 0.7% by GPT-4), existential theres (EX: 0.92% vs. 0.13% by GPT-3.5 and 0.09% by GPT-4), coordinating conjunctions (CC: 0.76% vs. 0.02% by GPT-4 and 0.11% by GPT-3.5), foreign words (FW: 0.45% vs. 0.03% by GPT-3.5 and 0.05% by GPT-4), list markers (LS: 0.55% vs. 0% by GPT-3.5 and 0% by GPT-4), wh-pronouns (WP: 0.58% vs. 0% by GPT-3.5 and 0.01% by GPT-4), wh-determiners (WDT: 1.00% vs. 0% by GPT-3.5 and 0.01% by GPT-4), and interjections (UH: 0.53% vs. 0.06% by GPT-3.5 and 0.08% by GPT-4). More broadly speaking, when non-negligible use of POSs is considered (i.e., over 0.25% prevalence), Bard presents conspicuously diverse use patterns (88.5%) compared to GPT-3.5 and GPT-4 (48.6% and 60.0%, respectively).

Overall, notable differences are observed in the POS distributions associated with the examined LLMs. The differences seem to span over both highly frequent POSs (i.e., noun) and low-frequency POSs (i.e., existential there). Most notably, significant differences are observed in the presence of low-frequency POSs with Bard typically using them more often.

3.3. Dependency Parsing

Next, we examine the distribution of key dependencies in the provided responses, as summarized in Figure 2 (the complete list, along with their abbreviations and complete results, are provided in Appendix A). Each dependency represents a directed syntactic relation between two tokens, where one token functions as the syntactic head and the other as its dependent. For example, in a dependency relation such as nominal subject, the subject token is linked to the predicate it depends on. Since the parser treats punctuation as tokens, punctuation marks are included in the resulting dependency graph and may be assigned a punctuation dependency relation. The results are summarized in Figure 4. Statistically, GPT-3.5 and GPT-4 do not differ significantly in terms of dependencies (p = 0.13) while BARD does differ from both at p < 0.05.

Top 30 dependency relations comparison between GPT-3.5, GPT-4, and Bard
Figure 2. Top 30 dependency relations comparison between GPT-3.5, GPT-4, and Bard.

Given the wide variety of dependency types, we discuss the prominent differences only for dependencies with at least 1% prevalence (for at least one of the examined LLMs). Starting with the three most prevalent dependency types in the data, punctuation, determiner, and preposition, we see that Bard uses them more often than GPT-3.5 and GPT-4 (punct: 10.5% vs. 9.62%, 8.56%; det: 10.06% vs. 8.84%, 8.46%; prep: 10.78% vs. 8.73%, 8.95%). Other noteworthy differences include Bard’s higher use of nominal subjects (nsubj: 9.17% vs. 7.72%, 7.23%), direct objects (dobj: 6.44% vs. 5.13%, 5.44%), auxiliary (aux: 8.12% vs. 5.12%, 4.96%) and coordinating conjunctions (cc: 6.23% vs. 4.18%, 3.93%), and lower use of adjectival modifier (amod: 3.72% vs. 5.13%, 5.44%), conjunct (conj: 3.07% vs. 4.87%, 5.2%), root dependencies (ROOT: 2.06% vs. 3.92%, 4.19%), adverbial clause modifier (advcl: 0.23% vs. 2.24% and 2.33%), adjectival complement (acomp: 0.67% vs. 1.85%, 2.07%), clausal complement (ccomp: 0.08% vs. 1.79%, 1.66%), possession modifiers (poss: 0.72% vs. 1.62%, 2.45%), passive auxiliaries (auxpass: 0% vs. 1.39%, 0.99%) and attributions (attr: 0% vs. 1.11%, 1.27%).

Overall, while no significant differences are observed between GPT-3.5 and GPT- 4, Bard seems to significantly differ from both with a distinctively higher use of highly frequent dependencies (i.e., punctuation, determination, preposition, direct object, auxiliary verb, etc) and a lower use of other dependencies (i.e., adverbial clause modifier, adjectival complement, clausal complement, attributions, etc).

3.4. Sentiment Analysis

Considering sentiment, the three LLMs do not present statistically significant differences p = 0.284, as presented in Figure 3. Therefore, sentiment should not be interpreted as an independently distinctive marker for differentiating between GPT-3.5, GPT-4, and Bard. Rather, the sentiment distributions suggest that the three examined LLMs express broadly similar affective orientations in their responses. All three LLMs tend to produce positive sentiment (GPT-3.5: 53.8%; GPT-4: 53.3%; and Bard: 53.22%) with relatively few cases of neutral sentiment (GPT-3.5: 2.96%; GPT-4: 3.75%; and Bard: 2.89%).

3.5. Baseline and Ablation Analysis

In order to further examine which linguistic feature families contribute to the LLM-attribution performance, we conducted additional baseline and ablation analyses. First, we considered a majority-class baseline, which always predicts the most frequent class. Since the attribution task is balanced across the three examined LLMs, this baseline obtains an accuracy of approximately 0.33. Second, we trained separate models using each feature family independently: vocabulary features, POS features, dependency features, and sentiment features. Finally, we conducted leave-one-feature-family-out ablations, in which each feature group was removed from the full linguistic feature set in turn.

Sentiment distribution over the responses generated by GPT-3.5, GPT-4, and Bard
Figure 3. Sentiment distribution over the responses generated by GPT-3.5, GPT-4, and Bard.

The results are reported in Table 2. Overall, the results indicate that attribution performance is not driven by a single linguistic marker alone. Rather, vocabulary, POS, and dependency features all contribute to the model’s performance, whereas sentiment alone provides limited discriminatory power. The best performance is obtained when all feature families are combined, suggesting that the attribution model relies on a multivariate linguistic profile of the generated text.

Table 2. Baseline and ablation results for the LLM-attribution task. The results are shown as the mean of k = 5 folds.
Feature SettingAccuracyF1 Score
Majority-class baseline0.330.17
Vocabulary only0.730.72
POS only0.790.78
Dependency only0.760.75
Sentiment only0.360.31
All features except vocabulary0.840.83
All features except POS0.820.81
All features except dependency0.830.82
All features except sentiment0.870.86
All features0.880.87

3.6. LLM Attribution

Here, we consider the task of LLM attribution. That is, given a text, we wish to classify it to its LLM source. Recall that in this study, we are primarily interested in investigating the linguistic differences between LLMs and their potential use as indicators for LLM attribution. As such, we opt for an off-the-shelf XGboost model that receives the above linguistic profile of a given text as input (i.e., vocabulary, POS distribution, dependency distribution, and sentiment) and classifies it into one of the three LLMs (i.e., its origin).

We used a grouped k-fold cross-validation approach with k = 5. The grouping unit was the original HC3 input prompt. That is, for each sampled input qi, the corresponding responses generated by GPT-3.5, GPT-4, and Bard were always assigned to the same fold. Therefore, responses originating from the same prompt could not appear in both the training and test sets. The split was additionally stratified by dataset, such that in each fold the relative portion of train and test instances from each of the five HC3 datasets was identical. Concretely, each fold used approximately 80% of the original prompts for training and 20% for testing, with all model-specific responses to a given prompt kept together. Note that, presumably, more advanced and fine-tuned models could be developed for the task of LLM attribution; however, these are outside the focus of our analysis.

Our results show a high accuracy of 0.88 with F1 score of 0.87. Table 3 presents the classification performance of the XGboost model for each LLM and the average performance. The results are shown as the mean of k = 5 folds.

Table 3. Classification performance of the linguistic-based XGboost model for each LLM and the average performance. The results are shown as the mean of k = 5 folds.
PredictionRecallF1 ScoreAccuracySupport
GPT-3.50.900.820.860.871000
GPT-40.820.840.830.851000
Bard0.940.890.910.921000
Average0.890.850.870.883000

Feature importance is extracted using the information gain method and is reported in Figure 4. Note that linguistic features from all examined aspects (i.e., vocabulary, POS, dependency, and sentiment) are present in the top ten most important features of the model; however, this should be understood as their relative contribution to the multivariate classifier rather than as evidence that each feature type is independently distinctive across the three LLMs. These are led by the noun (NN) and proper noun (NNP) POSs, positive sentiment, punctuation dependency (punct), and the vocabulary’s density and word count. Importantly, feature importance in the XGBoost model should not be interpreted in the same way as the univariate statistical comparisons reported above. A feature may contribute to the classifier in tandem with other features, even when it does not show a statistically significant difference when examined independently. Accordingly, the appearance of sentiment-related features among the top-ranked features does not imply that sentiment is, by itself, a distinctive marker of a specific LLM. Instead, it indicates that sentiment provides some information gain within the multivariate classification model, in interaction with the broader linguistic profile of the text.

Additional baseline and ablation analyses are reported in Section 3.5. These analyses align with the above results and further indicate that the attribution performance is driven primarily by the combination of vocabulary, POS, and dependency features, rather than solely by sentiment or by any single feature family.

Feature importance of the top 10 most important features of the XGboost model
Figure 4. Feature importance of the top 10 most important features of the XGboost model.

4. Discussion

In this study, we analyzed and compared the linguistic characteristics of three of the most popular LLMs today: GPT-3.5, GPT-4, and Bard. Our results combine to suggest that, just like human authors, LLMs are linguistically distinguishable. Specifically, the results show that different LLMs present different linguistic markers, especially in terms of vocabulary, POS, and dependencies. In contrast, sentiment does not significantly differ across the three examined LLMs when analyzed independently, and should therefore be viewed only as a feature that may contribute to attribution in combination with other linguistic markers. These and similar differences are shown to be effectively used for LLM attribution—i.e., the identification of which LLM generated a given text, with a remarkably high accuracy.

This key result has important theoretical and practical implications. From a theoretical perspective, this result seems to affirm the potential of LLMs to reproduce some of the diversity inherent in human language expression. That is, different LLMs tend to display different linguistic styles, akin to how human authors do. Furthermore, the observed differences could advance our understanding of the potential impact of various design and training decisions on the linguistic style presented by an LLM. We plan to conduct such an inquiry in future work. From a practical perspective, the identified differences could aid in model comparison and evaluation, and guide the selection and development of LLMs for specific purposes and applications [38–40]. Moreover, we believe that the results should be taken into consideration for the important task of LLM detection [41,42]. The present findings suggest that LLM detection should not be viewed solely as a binary distinction between human-generated and machine-generated text [43]. Instead, LLM-generated texts may themselves constitute a heterogeneous class, with different models exhibiting different linguistic profiles. This observation may have important implications, as many detection models are trained and evaluated on texts generated by a limited set of models and may therefore learn markers specific to those models rather than to LLM-generated text more generally. Specifically, a detector trained primarily on outputs from one LLM may show reduced reliability when applied to outputs generated by another LLM, especially when the models differ in vocabulary, POS distribution, dependency structure, or other stylistic properties. Therefore, we speculate that future LLM-detection systems could benefit from incorporating model-specific linguistic markers, both to improve robustness across different LLMs and to provide more interpretable evidence for detection decisions. From this perspective, LLM attribution and LLM detection are complementary tasks: attribution can reveal which linguistic markers distinguish among different generators, while detection can leverage this information to construct broader, more reliable human-versus-LLM classifiers. At the same time, it is important to note that our study should not be interpreted as proposing a complete LLM-detection framework; rather, it highlights a source of variation that detection systems could capitalize on when targeting cross-model generalization.

This study is not without limitations. First, our analysis focuses on a diverse, yet limited, cohort of datasets. As such, an investigation into different datasets may result in slightly different outcomes, depending on the nature of the inputs provided to the LLM. Second, the linguistic style is collected for the zero-shot case where the LLMs do not have any context. Thus, our study focuses on the most simplistic use-case of these advanced tools where context does not play a role. Future works may replicate our analysis with different datasets and contexts, providing a wider perspective on the matter. Additionally, our analysis focuses on the English language. Studying the differences between LLMs across languages can be valuable to a wider range of applications and use cases.

Author Contributions: Conceptualization, T.L. and A.R.; methodology, T.L. and A.R.; software, T.L. and A.R.; validation, T.L. and A.R.; formal analysis, T.L. and A.R.; investigation, T.L. and A.R.; data curation, T.L. and A.R.; writing—original draft, T.L. and A.R.; writing—review & editing, T.L.; visualization, T.L. and A.R.; supervision, T.L.; and project administration, A.R. All authors have read and agreed to the published version of the manuscript.

Funding: This research received no external funding.

Data Availability Statement: The data presented in this study are openly available in [Github].

Conflicts of Interest: The authors declare no conflicts of interest

Appendix A

Appendix A.1. Prompting Protocol Since prompting may affect the linguistic style of LLM-generated text, we report here the exact prompting protocol used to generate the additional responses analyzed in this study. For each input qi sampled from HC3, the same prompt template was used for GPT-4 and Bard:

Please answer the following question: {qi}

That is, the original HC3 input was submitted verbatim to the model, without additional task-specific instructions, style instructions, examples, or model-specific modifications. The GPT-3.5 responses were taken from the original HC3 corpus, which pairs the same inputs with GPT-3.5-generated answers. Thus, all three LLMs were compared on responses to the same underlying inputs.

Appendix A.2. Additional Results

Table A1. Sentiment distribution.
LLMNegativeNeutralPositive
GPT-3.543.242.9653.80
GPT-442.953.7553.30
Bard43.892.8953.22
POSSymbolGPT-3.5GPT-4Bard
noun singularNN30.3628.7326.15
noun pluralNNS16.1913.5112.67
adjectiveJJ15.1214.1512.75
proper noun, singularNNP8.5511.439.91
verb, gerund/present participleVBG4.663.933.84
verb, singular, present, non-3dVBP4.614.453.83
adverbRB3.664.874.15
verb, base formVB3.043.422.96
preposition/subordinating conjunctionIN2.272.522.49
verb, 3rd person singular presentVBZ2.122.111.77
verb, past participleVBN1.81.812.43
verb, past tenseVBD1.791.882.74
modalMD1.581.991.45
determinerDT1.551.91.45
adjective, comparativeJJR0.720.680.89
personal pronounPRP0.690.71.7
cardinal digitCD0.210.350.18
adjective, superlativeJJS0.190.260.82
proper noun, pluralNNPS0.170.290.26
adverb, comparativeRBR0.150.171.25
existential thereEX0.130.090.92
possessive pronounPRP$0.10.280
wh-abverbWRB0.10.080
to go ‘to’TO0.080.050.65
interjectionUH0.060.080.53
particleRP0.040.060
foreign wordFW0.030.050.45
coordinating conjunctionCC0.020.110.76
possessive endingRBS0.020.010.27
possessive wh-pronounWP$00.010
wh-determinerWDT00.011
wh-pronounWP00.010.58
predeterminerPDT000.59
list markerLS000.55
DependencySymbolGPT-3.5GPT-4Bard
punctuationpunct9.628.5610.5
determinerdet8.848.4610.06
Prepositional modifierprep8.738.9510.78
Preposition objectpobj8.367.957.77
Nominal subjectnsubj7.727.239.17
Adjectival modifieramod5.725.333.72
Direct Objectdobj5.135.446.44
Auxiliary verbaux5.124.968.12
Conjunctconj4.875.23.07
Coordinating conjunctioncc4.183.936.23
Adverbial modifieradvmod4.084.13.49
RootROOT3.924.192.06
Compound wordcompound3.173.053.69
Adverbial clause modifieradvcl2.242.330.23
Open clausal complementxcomp2.012.022.3
Adjectival complementacomp1.852.070.67
Clausal complementccomp1.791.660.08
Markermark1.712.472.47
Relative clause modifierrelcl1.661.61.5
Possession modifierposs1.622.450.72
Auxiliary verb (passive)auxpass1.390.990
Attributeattr1.111.270
Nominal subject (passive)nsubjpass0.990.931.1
Unclassified dependentdep0.90.881.02
Preposition complementpcomp0.840.761.04
Clausal modifier of nounacl0.640.640.71
Numbernummod0.640.560.68
Verb particleprt0.540.570.54
Expletiveexpl0.370.370.5
Object predicateoprd0.340.290.37
Negation modifierneg0.320.290.38
Appositional modifierappos0.190.190.2
Clausal subject (passive)csubjpass0.190.050.19
Noun phrase as adverbial modifiernpadvmod0.130.10.14
Dativedative0.080.080.06
Meta datameta0.050.010
Agent (passive)agent000
Modifier of nominalnmod000
Parataxisparataxis00.060
Case markercase000
Clausal subjectcsubj00.010

Article notes

Publication history
Received 15 April 2026 · Accepted 8 May 2026 · Published 14 May 2026
Keywords
  • Large Language Models
  • writing style
  • linguistic analysis

References

  1. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 2022, 35, 27730–27744.
  2. Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. Gpt-4 technical report. arXiv 2023, arXiv:2303.08774. doi:10.48550/arXiv.2303.08774
  3. McGowan, A.; Gui, Y.; Dobbs, M.; Shuster, S.; Cotter, M.; Selloni, A.; Goodman, M.; Srivastava, A.; Cecchi, G.A.; Corcoran, C.M. ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Res. 2023, 326, 115334. doi:10.1016/j.psychres.2023.115334
  4. Zhao, W.X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. A survey of large language models. arXiv 2023, arXiv:2303.18223. doi:10.1007/s11704-026-60308-3
  5. Ben-Zion, Z.; Simon, G.; Lazebnik, T. Can LLMs Get High? A Dual-Metric Framework for Evaluating Psychedelic Simulation and Safety in Large Language Models. Res. Sq. 2026.
  6. Taskin, B.; Xie, W.; Lazebnik, T. Knowledge integration for physics-informed symbolic regression using pre-trained large language models. Sci. Rep. 2026, 16, 1614. doi:10.1038/s41598-026-35327-6
  7. Lazebnik, T.; Shami, L. Investigating tax evasion emergence using dual large language model and deep reinforcement learning powered agent-based simulation. arXiv 2025, arXiv:2501.18177. doi:10.48550/arXiv.2501.18177
  8. Sadasivan, V.S.; Kumar, A.; Balasubramanian, S.; Wang, W.; Feizi, S. Can AI-Generated Text be Reliably Detected? arXiv 2023, arXiv:2303.11156.
  9. Casal, J.E.; Kessler, M. Can linguists distinguish between ChatGPT/AI and human writing?: A study of research ethics and academic publishing. Res. Methods Appl. Linguist. 2023, 2, 100068. doi:10.1016/j.rmal.2023.100068
  10. Tang, R.; Chuang, Y.N.; Hu, X. The Science of Detecting LLM-Generated Texts. arXiv 2023, arXiv:2303.07205. doi:10.1145/3624725
  11. Deng, Z.; Gao, H.; Miao, Y.; Zhang, H. Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model. arXiv 2023, arXiv:2305.16617. doi:10.48550/arXiv.2305.16617
  12. Koike, R.; Kaneko, M.; Okazaki, N. OUTFOX: LLM-generated Essay Detection through In-context Learning with Adversarially Generated Examples. arXiv 2023, arXiv:2307.11729. doi:10.1609/aaai.v38i19.30120
  13. Herbold, S.; Hautli-Janisz, A.; Heuer, U.; Kikteva, Z.; Trautsch, A. A large-scale comparison of human-written versus ChatGPT-generated essays. Sci. Rep. 2023, 13, 18617.
  14. Muñoz-Ortiz, A.; Gómez-Rodríguez, C.; Vilares, D. Contrasting linguistic patterns in human and llm-generated text. arXiv 2023, arXiv:2308.09067. doi:10.1007/s10462-024-10903-2
  15. Bhattacharjee, A.; Kumarage, T.; Moraffah, R.; Liu, H. ConDA: Contrastive Domain Adaptation for AI-generated Text Detection. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 598–610.
  16. Su, J.; Zhuo, T.Y.; Wang, D.; Nakov, P. DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine- Generated Text. arXiv 2023, arXiv:2306.05540.
  17. Giorgi, S.; Markowitz, D.M.; Soni, N.; Varadarajan, V.; Mangalik, S.; Schwartz, H.A. I slept like a baby: Using human traits to characterize deceptive ChatGPT and human text. In Proceedings of the International workshop on implicit author characterization from texts for search and retrieval (IACT’23), Taipei, Taiwan, 27 July 2023 .
  18. Desaire, H.; Chua, A.E.; Isom, M.; Jarosova, R.; Hua, D. Distinguishing academic science writing from humans or ChatGPT with over 99% accuracy using off-the-shelf machine learning tools. Cell Rep. Phys. Sci. 2023, 4, 101426. doi:10.1016/j.xcrp.2023.101426
  19. Wu, J.; Yang, S.; Zhan, R.; Yuan, Y.; Wong, D.F.; Chao, L.S. A survey on llm-gernerated text detection: Necessity, methods, and future directions. arXiv 2023, arXiv:2310.14724.
  20. Savchenko, E.; Lazebnik, T. Computer aided functional style identification and correction in modern Russian texts. J. Data Inf. Manag. 2022, 4, 25–32. doi:10.1007/s42488-021-00062-2
  21. McNamara, D.S.; Crossley, S.A.; McCarthy, P.M. The linguistic features of writing quality. Writ. Commun. 2010, 27, 57–86. doi:10.1177/0741088309351547
  22. Snow, E.L.; Allen, L.K.; Jacovina, M.E.; Perret, C.A.; McNamara, D.S. You’ve Got Style: Detecting Writing Flexibility across Time. In Proceedings of the Fifth International Conference on Learning Analytics and Knowledge; Association for Computing Machinery: New York, NY, USA, 2015; pp. 194–202.
  23. Matsuda, P.K.; Tardy, C.M. Voice in academic writing: The rhetorical construction of author identity in blind manuscript review. Engl. Specif. Purp. 2007, 26, 235–249. doi:10.1016/j.esp.2006.10.001
  24. Lample, G.; Subramanian, S.; Smith, E.; Denoyer, L.; Ranzato, M.; Boureau, Y. Multiple-Attribute Text Rewriting. In Proceedings of the International Conference on Learning Representations, New Orleans, LO, USA, 6–9 May 2019.
  25. Hartley, J.; Pennebaker, J.W.; Fox, C. Using New Technology to Assess the Academic Writing Styles of Male and Female Pairs and Individuals. J. Tech. Writ. Commun. 2003, 33, 243–261. doi:10.2190/9VPN-RRX9-G0UF-CJ5X
  26. Hartley, J.; Howe, M.; McKeachie, W. Writing through time: Longitudinal studies of the effects of new technology on writing. Br. J. Educ. Technol. 2001, 32, 141–151. doi:10.1111/1467-8535.00185
  27. Lagutina, K.; Lagutina, N.; Boychuk, E.; Vorontsova, I.; Shliakhtina, E.; Belyaeva, O.; Paramonov, I.; Demidov, P.G. A Survey on Stylometric Text Features. In Proceedings of the 2019 25th Conference of Open Innovations Association (FRUCT), Helsinki, Finland, 5–8 November 2019; pp. 184–195.
  28. Guo, B.; Zhang, X.; Wang, Z.; Jiang, M.; Nie, J.; Ding, Y.; Yue, J.; Wu, Y. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. arXiv 2023, arXiv:2301.07597. doi:10.48550/arXiv.2301.07597
  29. Maia, M.; Handschuh, S.; Freitas, A.; Davis, B.; McDermott, R.; Zarrouk, M.; Balahur, A. WWW’18 open challenge: Financial opinion mining and question answering. In Proceedings of the Web Conference 2018, Lyon, France, 23–27 April 2018; pp. 1941–1942.
  30. Zeng, G.; Yang, W.; Ju, Z.; Yang, Y.; Wang, S.; Zhang, R.; Zhou, M.; Zeng, J.; Dong, X.; Zhang, R.; et al. MedDialog: Large-scale Medical Dialogue Datasets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP); Webber, B., Cohn, T., He, Y., Liu, Y., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 9241–9250. doi:10.18653/v1/2020.emnlp-main.743
  31. Fan, A.; Jernite, Y.; Perez, E.; Grangier, D.; Weston, J.; Auli, M. ELI5: Long Form Question Answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics; Korhonen, A., Traum, D., Màrquez, L., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 3558–3567. doi:10.18653/v1/P19-1346
  32. Yang, Y.; Yih, W.t.; Meek, C. WikiQA: A Challenge Dataset for Open-Domain Question Answering. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing; Màrquez, L., Callison-Burch, C., Su, J., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2015; pp. 2013–2018. doi:10.18653/v1/D15-1237
  33. Nadkarni, P.M.; Ohno-Machado, L.; Chapman, W.W. Natural language processing: An introduction. J. Am. Med. Inform. Assoc. 2011, 18, 544–551. doi:10.1136/amiajnl-2011-000464
  34. Lehmann, C. The nature of parts of speech. Sprachtypol. Universalienforschung 2013, 66, 141–177.
  35. Di Caro, L.; Grella, M. Sentiment analysis via dependency parsing. Comput. Stand. Interfaces 2013, 35, 442–453. doi:10.1016/j.csi.2012.10.005
  36. Wankhade, M.; Rao, A.C.S.; Kulkarni, C. A survey on sentiment analysis methods, applications, and challenges. Artif. Intell. Rev. 2022, 55, 5731–5780. doi:10.1007/s10462-022-10144-1
  37. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794.
  38. Ghosal, S.S.; Chakraborty, S.; Geiping, J.; Huang, F.; Manocha, D.; Bedi, A. A Survey on the Possibilities & Impossibilities of AI-generated Text Detection. Trans. Mach. Learn. Res. 2023.
  39. Verma, V.; Fleisig, E.; Tomlin, N.; Klein, D. Ghostbuster: Detecting Text Ghostwritten by Large Language Models. arXiv 2023, arXiv:2305.15047.
  40. Macko, D.; Moro, R.; Uchendu, A.; Srba, I.; Lucas, J.S.; Yamashita, M.; Tripto, N.I.; Lee, D.; Simko, J.; Bielikova, M. Authorship Obfuscation in Multilingual Machine-Generated Text Detection. arXiv 2024, arXiv:2401.07867. doi:10.48550/arXiv.2401.07867
  41. Glukhov, D.; Shumailov, I.; Gal, Y.; Papernot, N.; Papyan, V. LLM Censorship: A Machine Learning Challenge or a Computer Security Problem? arXiv 2023, arXiv:2307.10719. doi:10.48550/arXiv.2307.10719
  42. Wu, J.; Hooi, B. Fake News in Sheep’s Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks. arXiv 2023, arXiv:2310.10830.
  43. Wang, Y.; Mansurov, J.; Ivanov, P.; Su, J.; Shelmanov, A.; Tsvigun, A.; Whitehouse, C.; Afzal, O.M.; Mahmoud, T.; Aji, A.F.; et al. M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection. arXiv 2023, arXiv:2305.14902.

This page reproduces the article Rosenfeld et al. (2026), Mathematics, doi:10.3390/math14101683, with the permission of the publisher. Text, tables and figures were extracted from the PDF and the layout adapted for the web; the PDF is the version of record.

Cite this paper

APA

Rosenfeld, A., & Lazebnik, T. (2026). Whose LLM Is It Anyway? Linguistic Comparison and LLM Attribution for GPT-3.5, GPT-4 and Bard. Mathematics, 14, 1683. https://doi.org/10.3390/math14101683

BibTeX

@article{rosenfeld2026whose,
  title = {Whose LLM Is It Anyway? Linguistic Comparison and LLM Attribution for GPT-3.5, GPT-4 and Bard},
  author = {Rosenfeld, Ariel and Lazebnik, Teddy},
  journal = {Mathematics},
  volume = {14},
  pages = {1683},
  year = {2026},
  doi = {10.3390/math14101683}
}