🌱 agrobox

Statistical analysis, visualization and decision support for agricultural experiments

CRAN status CRAN downloads R-CMD-check


🌾 Why agrobox?

Agricultural experiments can be carefully designed, rigorously conducted, and full of valuable information.

But after collecting the data, researchers often face another challenge:

How do I turn my experimental data into a statistical result β€” and, more importantly, into a decision?

The statistical workflow can quickly become complicated:

Data β†’ Choose the model β†’ Check assumptions β†’ ANOVA β†’ Post-hoc β†’ Interpret β†’ Figure β†’ Decision

For researchers who are not specialized in R or statistics, this becomes a real barrier.

agrobox was created to reduce that barrier.

The goal is not to replace experimental design or statistical thinking. The goal is to make the analytical workflow:

⚠️ agrobox does not rescue a poorly designed experiment.
It helps you get more efficiently from a well-designed experiment to its analysis and interpretation.


πŸš€ What is agrobox?

agrobox is an R package designed for statistical analysis, visualization, and reporting of agricultural and agroindustrial experiments.

It provides automated workflows for:


πŸ“¦ Installation

Install the stable version from CRAN:

install.packages("agrobox")

Then load the package:

library(agrobox)

🧠 Core idea: agrobox()

The main function organizes the entire statistical workflow around your experimental factors and response variable:

resultado <- agrobox(
  data     = datos,
  factor   = "tratamiento",
  variable = "rendimiento"
)

resultado

The function automatically evaluates the statistical assumptions and selects the appropriate analysis path. The goal is simple: place the grouping letters through a viable and defensible statistical route.

Clusters (one panel per group combination)
        ↓
Sufficient data? (β‰₯ 2 treatments, β‰₯ 3 obs. per treatment)
        ↓
Shapiro-Wilk (normality of residuals)
        ↓
Fligner-Killeen (homogeneity of variances)
        ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ A  Normal +          β”‚  β”‚ B  Normal +          β”‚  β”‚ C  Not normal +      β”‚  β”‚ D  Not normal +      β”‚
β”‚    homogeneous       β”‚  β”‚    heteroscedastic   β”‚  β”‚    homogeneous       β”‚  β”‚    heteroscedastic   β”‚
β”‚                      β”‚  β”‚                      β”‚  β”‚                      β”‚  β”‚                      β”‚
β”‚       ANOVA          β”‚  β”‚     Welch ANOVA      β”‚  β”‚  Kruskal-Wallis      β”‚  β”‚  Same as C + note    β”‚
β”‚         ↓            β”‚  β”‚          ↓           β”‚  β”‚  (or Friedman with   β”‚  β”‚  on heterogeneous    β”‚
β”‚  Tukey / Duncan      β”‚  β”‚    Games-Howell      β”‚  β”‚   blocks) / Dunn     β”‚  β”‚  variances           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Block and second factor are then handled within the selected route. If a post-hoc test cannot be computed, the next viable route is tried, and each figure includes a note explaining which route was used and why.

Since version 0.4.0 the route is selected automatically; the var.equal argument is kept only for backward compatibility.

The researcher does not need to manually reproduce every step for every variable.


πŸ”¬ Statistical workflow

Parametric analysis - One-way and two-way ANOVA - Tukey HSD - Duncan Multiple Range Test

Robust analysis (when variance homogeneity is not supported) - Welch ANOVA - Games-Howell post-hoc test

Non-parametric analysis (when residual normality is not supported) - Kruskal-Wallis with agricolae letters or Dunn post-hoc test (np_test) - Friedman test for blocked designs (RCBD) - P-value adjustment selectable with p.adj (default Bonferroni)

Diagnostics - Shapiro-Wilk test for residual normality - Fligner-Killeen test for homogeneity of variances

Additional output - Coefficient of variation (CV) - Statistical power - Means, grouping letters, and significance annotations - Statistical route used in each panel ($stats: route, method, p-value and notes)


πŸ“Š Visualization

agrobox() returns ggplot2-based figures ready for publication, including:

The output can be further customized using the full ggplot2 ecosystem.


🧩 Multiple factors and experimental structures

Agricultural experiments frequently involve more than one factor (e.g., Variety Γ— Treatment, Treatment Γ— Location).

agrobox() supports one and two experimental factors, with options for:


🧠 agrosintesis() β€” From numbers to decisions

A single experiment rarely measures only one variable. You might record yield, fruit weight, firmness, color, soluble solids, acidity, incidence, severity β€” and more.

Running each analysis independently produces a lot of output without necessarily making the experiment easier to understand.

agrosintesis() solves this.

It applies the agrobox() workflow to multiple response variables simultaneously and consolidates the results into a structured, decision-oriented synthesis:

resultado <- agrosintesis(
  data      = datos,
  variables = c("rendimiento", "peso_fruto", "firmeza", "solidos_solubles")
)

Instead of:

Variable 1 β†’ analysis
Variable 2 β†’ analysis
Variable 3 β†’ analysis

You get:

EXPERIMENT
    ↓
Variable 1 + Variable 2 + Variable 3
    ↓             ↓             ↓
 Analysis      Analysis      Analysis
         \        |        /
          agrosintesis()
                ↓
           SYNTHESIS
                ↓
           DECISION

Statistical analysis should help you understand the experiment, not just produce more numbers.


πŸ“‘ agrotabla() β€” Publication-ready tables

Export statistical results as high-resolution images suitable for reports, presentations, and scientific publications:

agrotabla(resultado)

πŸ“Š agroexcel() β€” Excel export

Export results directly to Excel, organized by variable and experimental cluster:

agroexcel(resultado)

Particularly useful when an experiment contains several variables β€” results are organized into worksheets, eliminating manual copy-paste from R to Excel.


πŸ§ͺ A typical workflow

library(agrobox)

# Single variable
resultado <- agrobox(
  data     = datos,
  factor   = "tratamiento",
  variable = "rendimiento"
)

resultado

# Multiple variables
resultado <- agrosintesis(
  data      = datos,
  variables = c("rendimiento", "peso", "firmeza", "calidad")
)

# Export
agroexcel(resultado)
agrotabla(resultado)

The complete pipeline:

EXPERIMENTAL DATA
       ↓
   agrobox()
       ↓
Statistical diagnostics β†’ Analysis β†’ Post-hoc β†’ Graphics
       ↓
 agrosintesis()
       ↓
    SYNTHESIS
       ↓
    DECISION
       ↓
Excel / Tables

⚠️ What agrobox does NOT do

agrobox simplifies the statistical workflow. It does not replace experimental design or statistical reasoning.

No package can compensate for:

Good statistics cannot rescue bad experimental design.
When the experiment has been designed correctly, agrobox makes the analysis more accessible and reproducible.


πŸ’‘ Help shape agrobox

agrobox is open source. Its development is driven by real agricultural problems.

If you find yourself thinking β€œI wish agrobox could do this…” β€” please tell me.

When reporting a bug, please include: your R version, your agrobox version, a reproducible example, the error message, and what you expected to happen.


🀝 Contributing

Contributions are welcome. You can help by:

Every contribution helps make statistical analysis more accessible to agricultural researchers.


πŸ“š Scientific methods

The procedures implemented in agrobox are based on established statistical methods:

Method Reference
Tukey HSD Tukey (1949)
Duncan Multiple Range Test Duncan (1955)
Welch ANOVA Welch (1951)
Games-Howell Games & Howell (1976)
Shapiro-Wilk Shapiro & Wilk (1965)
Fligner-Killeen Fligner & Killeen (1976)
Kruskal-Wallis Kruskal & Wallis (1952)
Dunn test Dunn (1964)
Friedman test Friedman (1937)
Statistical power Cohen (1988)

The package builds on the R ecosystem, including ggplot2 and agricolae.


🌱 Citation

If you use agrobox in your research, please cite the package:

citation("agrobox")

πŸ‘¨β€πŸ”¬ Author

Joaquin Alejandro Salinas Angeles
Agronomist & Agricultural Researcher

Interests: agricultural experimentation Β· statistical analysis Β· postharvest research Β· reproducible research Β· R programming


⭐ Support the project

If agrobox is useful to you:


🌾 From experiment to decision β€” let’s make agricultural statistics more accessible.